← Back
AI Engineer July 31, 2026 20m

Agents at Scale: Inside MiniMax's Model and the Infrastructure Behind It — Olive Song

Read full transcript 17 segments
  1. This is a discussion that I'm This is a discussion that I'm particularly excited about because particularly excited about because particularly excited about because the field is moving so fast and we have the field is moving so fast and we have the field is moving so fast and we have two people that that kind of have this two people that that kind of have this two people that that kind of have this unique vantage point on the field and unique vantage point on the field and unique vantage point on the field and what I want to do is kind of ask what I want to do is kind of ask what I want to do is kind of ask questions to see if we can learn from questions to see if we can learn from questions to see if we can learn from that. So I want to start off with that. So I want to start off with that. So I want to start off with intros, talk a little bit about your intros, talk a little bit about your intros, talk a little bit about your role, what you're thinking about, what role, what you're thinking about, what role, what you're thinking about, what you're working on. Maybe Dan if you can you're working on. Maybe Dan if you can you're working on. Maybe Dan if you can go first. go first. go first. >> Uh hey everyone. I'm Dan. I'm the VP of >> Uh hey everyone. I'm Dan. I'm the VP of >> Uh hey everyone. I'm Dan. I'm the VP of Kernels Together AI. Kernels Together AI. Kernels Together AI. I lead inference, GPU optimization, I lead inference, GPU optimization, I lead inference, GPU optimization, trying to figure out how to use GPUs trying to figure out how to use GPUs trying to figure out how to use GPUs most effectively to serve AI models. most effectively to serve AI models. most effectively to serve AI models. >> Yeah. So one of the things that I wanted >> Yeah. So one of the things that I wanted >> Yeah. So one of the things that I wanted to dive in with Dan about is like you to dive in with Dan about is like you to dive in with Dan about is like you model drops, what are what is everything model drops, what are what is everything model drops, what are what is everything that goes on behind the scenes to serve that goes on behind the scenes to serve that goes on behind the scenes to serve it so that everybody here there's a lot it so that everybody here there's a lot it so that everybody here there's a lot of builders here that can use it. Olive of builders here that can use it. Olive of builders here that can use it. Olive I want to throw it over to you to talk I want to throw it over to you to talk I want to throw it over to you to talk about your role and what you're focusing about your role and what you're focusing about your role and what you're focusing on. on. on. >> Yeah. I'm Olive and I am the research >> Yeah. I'm Olive and I am the research >> Yeah. I'm Olive and I am the research lead of RL at MiniMax and I am lead of RL at MiniMax and I am lead of RL at MiniMax and I am responsible for the final training of responsible for the final training of responsible for the final training of the model and the shipping of the model. the model and the shipping of the model. the model and the shipping of the model. So basically everything before the So basically everything before the So basically everything before the inference, right? inference, right? inference, right? >> Okay. >> Okay. >> Okay. Awesome. Um so maybe I wanted to start Awesome. Um so maybe I wanted to start Awesome. Um so maybe I wanted to start off this panel is focusing on open off this panel is focusing on open off this panel is focusing on open source. I wanted to start off with this source. I wanted to start off with this source. I wanted to start off with this is your strongest model yet, MiniMax M3.

  2. is your strongest model yet, MiniMax M3. is your strongest model yet, MiniMax M3. Why open source it? What's the kind of Why open source it? What's the kind of Why open source it? What's the kind of the the idea behind that as a company as the the idea behind that as a company as the the idea behind that as a company as you're releasing these models? you're releasing these models? you're releasing these models? >> We do believe that the open source >> We do believe that the open source >> We do believe that the open source community as a whole is very strong and community as a whole is very strong and community as a whole is very strong and powerful. powerful. powerful. While we open source the model, everyone While we open source the model, everyone While we open source the model, everyone can use it. So it aligns with our can use it. So it aligns with our can use it. So it aligns with our mission that we want to have mission that we want to have mission that we want to have intelligence with everyone and also intelligence with everyone and also intelligence with everyone and also different developers can contribute to different developers can contribute to different developers can contribute to the model through feedbacks, through the model through feedbacks, through the model through feedbacks, through their own PRs and we can build the their own PRs and we can build the their own PRs and we can build the models even stronger. And also like for models even stronger. And also like for models even stronger. And also like for example, example, example, Dan you you will be able to optimize on Dan you you will be able to optimize on Dan you you will be able to optimize on our open weight model and make it our open weight model and make it our open weight model and make it inference faster and then serve better inference faster and then serve better inference faster and then serve better for everyone. for everyone. for everyone. >> Yeah. Yeah, we're we're big believers in >> Yeah. Yeah, we're we're big believers in >> Yeah. Yeah, we're we're big believers in open source that together and open source that together and open source that together and yeah, I think when we we've been you yeah, I think when we we've been you yeah, I think when we we've been you know, following you guys for for a know, following you guys for for a know, following you guys for for a while. I think from way old old older while. I think from way old old older while. I think from way old old older MiniMax models. So seeing M3 and seeing MiniMax models. So seeing M3 and seeing MiniMax models. So seeing M3 and seeing how far it's come is is really how far it's come is is really how far it's come is is really impressive and really great. impressive and really great. impressive and really great. >> So I wanted to kind of pick on >> So I wanted to kind of pick on >> So I wanted to kind of pick on this a little bit more. Can you explain this a little bit more. Can you explain this a little bit more. Can you explain so we've got the model creators so we've got the model creators so we've got the model creators themselves MiniMax. We've got experts on themselves MiniMax. We've got experts on themselves MiniMax. We've got experts on the inference side of things. How did the inference side of things. How did the inference side of things. How did this partnership come to be? So they this partnership come to be? So they this partnership come to be? So they launch an open source model launch an open source model launch an open source model and we're now distributing it. I I and we're now distributing it. I I and we're now distributing it. I I checked this morning. We have the lion's checked this morning. We have the lion's checked this morning. We have the lion's share of share of share of token usage for MiniMax M3. How does token usage for MiniMax M3. How does token usage for MiniMax M3. How does this partnership come to be and how do this partnership come to be and how do this partnership come to be and how do we serve a model like this at scale?

  3. we serve a model like this at scale? we serve a model like this at scale? >> Yeah, yeah, great question. So at >> Yeah, yeah, great question. So at >> Yeah, yeah, great question. So at together, I think one of the things that together, I think one of the things that together, I think one of the things that that we're really interested in is how that we're really interested in is how that we're really interested in is how do you do you do you make intelligence abundant? So how do make intelligence abundant? So how do make intelligence abundant? So how do you get more tokens to more people to do you get more tokens to more people to do you get more tokens to more people to do more useful things and and and get get more useful things and and and get get more useful things and and and get get all these capabilities into more all these capabilities into more all these capabilities into more people's hands. So we follow all the people's hands. So we follow all the people's hands. So we follow all the open models very closely. open models very closely. open models very closely. I I don't remember when exactly we we I I don't remember when exactly we we I I don't remember when exactly we we started part. Oh, actually I think I do started part. Oh, actually I think I do started part. Oh, actually I think I do know this. We had a car event in Las know this. We had a car event in Las know this. We had a car event in Las Vegas sometime last year. And with Vegas sometime last year. And with Vegas sometime last year. And with someone from MiniMax came and there he someone from MiniMax came and there he someone from MiniMax came and there he was like, guys, you really got to serve was like, guys, you really got to serve was like, guys, you really got to serve our next model. It's going to be really our next model. It's going to be really our next model. It's going to be really really great. really great. really great. So I think from there we we started So I think from there we we started So I think from there we we started talking. We were serving MiniMax 2.5 talking. We were serving MiniMax 2.5 talking. We were serving MiniMax 2.5 and I think 2.7 for for a while and I think 2.7 for for a while and I think 2.7 for for a while and then when leading up to the launch and then when leading up to the launch and then when leading up to the launch of of M3, of of M3, of of M3, we were we were quite excited about it. we were we were quite excited about it. we were we were quite excited about it. I think we we're seeing the the usage I think we we're seeing the the usage I think we we're seeing the the usage and what people were doing with it was and what people were doing with it was and what people were doing with it was it was really quite exciting it was really quite exciting it was really quite exciting and and so from there and and so from there and and so from there that's really where we partner. We start that's really where we partner. We start that's really where we partner. We start working on the model. The architecture, working on the model. The architecture, working on the model. The architecture, optimizing it, optimizing it, optimizing it, figuring out you know, what's the best figuring out you know, what's the best figuring out you know, what's the best way to serve inference on it and and all way to serve inference on it and and all way to serve inference on it and and all those great pieces.

  4. those great pieces. those great pieces. >> Yeah. I wanted to actually get into more >> Yeah. I wanted to actually get into more >> Yeah. I wanted to actually get into more on the model side of things. And so as on the model side of things. And so as on the model side of things. And so as the creator of a model, as as somebody the creator of a model, as as somebody the creator of a model, as as somebody who's like post-training this thing, the who's like post-training this thing, the who's like post-training this thing, the model lands and all the builders that model lands and all the builders that model lands and all the builders that are here start using it. are here start using it. are here start using it. From your perspective, um what are the From your perspective, um what are the From your perspective, um what are the kind of the unique capabilities that you kind of the unique capabilities that you kind of the unique capabilities that you love to see people use it for? And what love to see people use it for? And what love to see people use it for? And what are maybe some of the hidden gems that are maybe some of the hidden gems that are maybe some of the hidden gems that you haven't You thought, "Oh, people you haven't You thought, "Oh, people you haven't You thought, "Oh, people would love to build this." but you would love to build this." but you would love to build this." but you haven't seen a little bit. Could you haven't seen a little bit. Could you haven't seen a little bit. Could you shed more light on that? shed more light on that? shed more light on that? >> Mhm. Mhm. So, MiniMax M3, which was >> Mhm. Mhm. So, MiniMax M3, which was >> Mhm. Mhm. So, MiniMax M3, which was different from the M2 series, was that different from the M2 series, was that different from the M2 series, was that it was actually multi-modal it was actually multi-modal it was actually multi-modal So, it not only understands text and it So, it not only understands text and it So, it not only understands text and it not only writes code, it also not only writes code, it also not only writes code, it also understands videos and images. So, we understands videos and images. So, we understands videos and images. So, we did see a lot of applications on a gen did see a lot of applications on a gen did see a lot of applications on a gen tech multi-modal agents, which is very tech multi-modal agents, which is very tech multi-modal agents, which is very cool. cool. cool. Um and I would say there are a couple Um and I would say there are a couple Um and I would say there are a couple that we can highlight, right? For that we can highlight, right? For that we can highlight, right? For example, computer uses. The model can is example, computer uses. The model can is example, computer uses. The model can is able to navigate through a computer and able to navigate through a computer and able to navigate through a computer and then do some pretty good creations with then do some pretty good creations with then do some pretty good creations with the tools that they they can utilize. the tools that they they can utilize. the tools that they they can utilize. And also, um you can develop games with And also, um you can develop games with And also, um you can develop games with the model.

  5. the model. the model. Um it's very fun that I I don't I think Um it's very fun that I I don't I think Um it's very fun that I I don't I think that's one of the hidden gems is that we that's one of the hidden gems is that we that's one of the hidden gems is that we actually worked on actually worked on actually worked on um game development. So, the model can um game development. So, the model can um game development. So, the model can help you develop a real cool games. help you develop a real cool games. help you develop a real cool games. Um yeah. Um yeah. Um yeah. >> Yeah, one >> Yeah, one >> Yeah, one So, when the when the blog dropped and So, when the when the blog dropped and So, when the when the blog dropped and then the paper dropped, one of the then the paper dropped, one of the then the paper dropped, one of the things I noticed was that you guys things I noticed was that you guys things I noticed was that you guys highlighted SVG bench, you guys highlighted SVG bench, you guys highlighted SVG bench, you guys highlighted kernel bench, you also highlighted kernel bench, you also highlighted kernel bench, you also touched on OS world. Can you talk more touched on OS world. Can you talk more touched on OS world. Can you talk more about Like, what does it take to about Like, what does it take to about Like, what does it take to post-train, especially for those post-train, especially for those post-train, especially for those particular domains? particular domains? particular domains? >> Mhm. Mhm. Uh I would say it's >> Mhm. Mhm. Uh I would say it's >> Mhm. Mhm. Uh I would say it's the very important thing is the data and the very important thing is the data and the very important thing is the data and how we define the problems. Um it could how we define the problems. Um it could how we define the problems. Um it could be very different from different tasks. be very different from different tasks. be very different from different tasks. For example, let's say the kernel one, For example, let's say the kernel one, For example, let's say the kernel one, right? The right? The right? The It would be very important to design the It would be very important to design the It would be very important to design the environments of the data so that we can environments of the data so that we can environments of the data so that we can deliberately train reinforcement deliberately train reinforcement deliberately train reinforcement learning in those very complex learning in those very complex learning in those very complex environments and let the model to environments and let the model to environments and let the model to optimize the kernels themselves and optimize the kernels themselves and optimize the kernels themselves and iteratively improve the performance. iteratively improve the performance. iteratively improve the performance. >> And one one aspect that I wanted to talk >> And one one aspect that I wanted to talk >> And one one aspect that I wanted to talk to you about on the kernel development to you about on the kernel development to you about on the kernel development side of things, um where are you seeing side of things, um where are you seeing side of things, um where are you seeing open models when it comes to kernel open models when it comes to kernel open models when it comes to kernel development? You recently released a development? You recently released a development? You recently released a benchmark specifically for this. So, I benchmark specifically for this. So, I benchmark specifically for this. So, I was wondering if you could talk on that was wondering if you could talk on that was wondering if you could talk on that a little bit as a little bit as a little bit as >> Yeah, yeah, it's a it's a great >> Yeah, yeah, it's a it's a great >> Yeah, yeah, it's a it's a great question. So, I think we're seeing all question. So, I think we're seeing all question. So, I think we're seeing all sorts of models of the closed frontier sorts of models of the closed frontier sorts of models of the closed frontier models and the open models get models and the open models get models and the open models get increasingly better at writing kernels.

  6. increasingly better at writing kernels. increasingly better at writing kernels. So, we use models all the time when we So, we use models all the time when we So, we use models all the time when we are developing kernels and and and are developing kernels and and and are developing kernels and and and writing the optimization frameworks. writing the optimization frameworks. writing the optimization frameworks. I think the the interesting thing that I think the the interesting thing that I think the the interesting thing that we are starting to look at is this we are starting to look at is this we are starting to look at is this benchmark that that we recently released benchmark that that we recently released benchmark that that we recently released called um parallel kernel bench. So, it called um parallel kernel bench. So, it called um parallel kernel bench. So, it actually has a bunch of unsolved actually has a bunch of unsolved actually has a bunch of unsolved problems in it. So, we went around problems in it. So, we went around problems in it. So, we went around surveyed all the different ways they can surveyed all the different ways they can surveyed all the different ways they can serve model inference. serve model inference. serve model inference. And one of the interesting things that And one of the interesting things that And one of the interesting things that we found is that there's a lot of things we found is that there's a lot of things we found is that there's a lot of things that we can think of that would actually that we can think of that would actually that we can think of that would actually speed models up that there don't exist speed models up that there don't exist speed models up that there don't exist good kernels for. good kernels for. good kernels for. So, one of the reasons that we put that So, one of the reasons that we put that So, one of the reasons that we put that benchmark out was, you know, one thing benchmark out was, you know, one thing benchmark out was, you know, one thing that people worry about is like bench that people worry about is like bench that people worry about is like bench maxing or or overfitting to particular maxing or or overfitting to particular maxing or or overfitting to particular benchmarks. benchmarks. benchmarks. One of our intentions with this One of our intentions with this One of our intentions with this benchmark was that if you overfit to it, benchmark was that if you overfit to it, benchmark was that if you overfit to it, that's great cuz we'll go take those that's great cuz we'll go take those that's great cuz we'll go take those kernels and use them to to to accelerate kernels and use them to to to accelerate kernels and use them to to to accelerate the the the the the the inference and the the the the inference and the the the the inference and the development. development. development. >> Yeah, this is a really interesting >> Yeah, this is a really interesting >> Yeah, this is a really interesting point. A lot of people have problems point. A lot of people have problems point. A lot of people have problems with bench maxing, but the way I think with bench maxing, but the way I think with bench maxing, but the way I think about it is if researchers like you put about it is if researchers like you put about it is if researchers like you put all the really useful benchmarks out and all the really useful benchmarks out and all the really useful benchmarks out and we bench max on all of them and we bench max on all of them and we bench max on all of them and everything is in distribution, then everything is in distribution, then everything is in distribution, then that's a perfect world, right? That's a that's a perfect world, right? That's a that's a perfect world, right? That's a very useful model that we can then use.

  7. very useful model that we can then use. very useful model that we can then use. Um Um Um Okay, cool. So, I wanted to touch on the Okay, cool. So, I wanted to touch on the Okay, cool. So, I wanted to touch on the inference side of things now as well. inference side of things now as well. inference side of things now as well. So, like a new model drops like this, So, like a new model drops like this, So, like a new model drops like this, what does it take Could you take me like what does it take Could you take me like what does it take Could you take me like behind the scenes at the inference behind the scenes at the inference behind the scenes at the inference stack? And what does it take to go from stack? And what does it take to go from stack? And what does it take to go from day zero launch and then optimizing it day zero launch and then optimizing it day zero launch and then optimizing it week over week, month over month? week over week, month over month? week over week, month over month? >> Yeah, yeah, great question. So, when we >> Yeah, yeah, great question. So, when we >> Yeah, yeah, great question. So, when we partner with someone like MiniMax, we partner with someone like MiniMax, we partner with someone like MiniMax, we will get some early model details. So, will get some early model details. So, will get some early model details. So, for M3, for example, there are things for M3, for example, there are things for M3, for example, there are things like the minimax sparse attention like the minimax sparse attention like the minimax sparse attention and and some of those choices that were and and some of those choices that were and and some of those choices that were a little bit different from any model a little bit different from any model a little bit different from any model that's out there. And I think if you that's out there. And I think if you that's out there. And I think if you look at any of the open models now, they look at any of the open models now, they look at any of the open models now, they are all quite different from each other are all quite different from each other are all quite different from each other in different ways. So, there's different in different ways. So, there's different in different ways. So, there's different attention, different MOE choices, attention, different MOE choices, attention, different MOE choices, differences in quantization and and all differences in quantization and and all differences in quantization and and all these pieces. So, as soon as we get these pieces. So, as soon as we get these pieces. So, as soon as we get those details, we start writing kernels, those details, we start writing kernels, those details, we start writing kernels, benchmarking, figuring out is there benchmarking, figuring out is there benchmarking, figuring out is there existing kernels that work for it? Do we existing kernels that work for it? Do we existing kernels that work for it? Do we need to modify something? Do we need to need to modify something? Do we need to need to modify something? Do we need to write something from scratch? write something from scratch? write something from scratch? And then so day zero, we're trying to And then so day zero, we're trying to And then so day zero, we're trying to think about things like quality. So, think about things like quality. So, think about things like quality. So, when this model launches, is it going to when this model launches, is it going to when this model launches, is it going to have the quality that we all expect? Are have the quality that we all expect? Are have the quality that we all expect? Are we going to be able to provide the right we going to be able to provide the right we going to be able to provide the right the the right user experience?

  8. the the right user experience? the the right user experience? And then from there, as soon as it And then from there, as soon as it And then from there, as soon as it launches that day zero, we have a long launches that day zero, we have a long launches that day zero, we have a long list of things that we know, hey, we list of things that we know, hey, we list of things that we know, hey, we have to do this with the KV cache, we have to do this with the KV cache, we have to do this with the KV cache, we have to do this with the attention have to do this with the attention have to do this with the attention kernels, we have to do kernels, we have to do kernels, we have to do we're going to look at this part of the we're going to look at this part of the we're going to look at this part of the quantization, and things like that. So, quantization, and things like that. So, quantization, and things like that. So, we we we we we have that list and then we start we we have that list and then we start we we have that list and then we start working on it and start optimizing over working on it and start optimizing over working on it and start optimizing over the course of weeks so that when you use the course of weeks so that when you use the course of weeks so that when you use these models, they actually get faster these models, they actually get faster these models, they actually get faster between day zero and day seven and day between day zero and day seven and day between day zero and day seven and day 14 and 14 and 14 and etc. etc. etc. >> Yeah, I was just talking to Ingrid >> Yeah, I was just talking to Ingrid >> Yeah, I was just talking to Ingrid actually yesterday and I and I asked actually yesterday and I and I asked actually yesterday and I and I asked her, have we been improving the her, have we been improving the her, have we been improving the performance of M3? performance of M3? performance of M3? And I meant over the last month and she And I meant over the last month and she And I meant over the last month and she said, oh, did you mean from last night? said, oh, did you mean from last night? said, oh, did you mean from last night? And this is the pace at which these guys And this is the pace at which these guys And this is the pace at which these guys work. So, work. So, work. So, it's very real. it's very real. it's very real. One aspect that I wanted to touch on One aspect that I wanted to touch on One aspect that I wanted to touch on with this is we're seeing the workloads with this is we're seeing the workloads with this is we're seeing the workloads shift. We're going from kind of shift. We're going from kind of shift. We're going from kind of predominantly chat workloads where you predominantly chat workloads where you predominantly chat workloads where you have turns coming in now to agentic have turns coming in now to agentic have turns coming in now to agentic workloads where you've got this thing workloads where you've got this thing workloads where you've got this thing sitting inside a harness and you're sitting inside a harness and you're sitting inside a harness and you're doing hundreds and hundreds of doing hundreds and hundreds of doing hundreds and hundreds of multi-turn tool calls. multi-turn tool calls. multi-turn tool calls. >> Yeah. >> Yeah. >> Yeah. >> Does that change the way you build the >> Does that change the way you build the >> Does that change the way you build the inference stack?

  9. inference stack? inference stack? >> Yeah, it definitely does. So, these >> Yeah, it definitely does. So, these >> Yeah, it definitely does. So, these agentic turn-based workloads, agentic turn-based workloads, agentic turn-based workloads, they go into everything from informing they go into everything from informing they go into everything from informing your KV cash, your KV cash, your KV cash, your prompting, your your prompting, your your prompting, your pieces like this. So, it it informs what pieces like this. So, it it informs what pieces like this. So, it it informs what part of the stack you want to go part of the stack you want to go part of the stack you want to go optimize optimize optimize because now I think when we're in the because now I think when we're in the because now I think when we're in the chat chat chat world, you have a system chat chat chat world, you have a system chat chat chat world, you have a system prompt of a few thousand, prompt of a few thousand, prompt of a few thousand, and then you just have the the chat and then you just have the the chat and then you just have the the chat logs. Now, with the coding base agentic logs. Now, with the coding base agentic logs. Now, with the coding base agentic workflows, you'll upload your whole code workflows, you'll upload your whole code workflows, you'll upload your whole code base to the model, and that's a very base to the model, and that's a very base to the model, and that's a very different optimization and routing and different optimization and routing and different optimization and routing and kernel challenge than than just the the kernel challenge than than just the the kernel challenge than than just the the chat base workload. So, yeah, we're we chat base workload. So, yeah, we're we chat base workload. So, yeah, we're we we follow these work codes very closely. we follow these work codes very closely. we follow these work codes very closely. It's really interesting to see how they It's really interesting to see how they It's really interesting to see how they evolve and and how to adapt the evolve and and how to adapt the evolve and and how to adapt the inference stack and the inference inference stack and the inference inference stack and the inference engines to to really take to really engines to to really take to really engines to to really take to really serve them well. serve them well. serve them well. >> Not only do you have >> Not only do you have >> Not only do you have agentic workloads, but you've also got agentic workloads, but you've also got agentic workloads, but you've also got multimodal workloads in there. So, what multimodal workloads in there. So, what multimodal workloads in there. So, what I like to do often with these coding I like to do often with these coding I like to do often with these coding agents is agents is agents is get them to optimize a web app and then get them to optimize a web app and then get them to optimize a web app and then get it to use it and then do a feedback get it to use it and then do a feedback get it to use it and then do a feedback loop. So, one thing that I wanted to loop. So, one thing that I wanted to loop. So, one thing that I wanted to come to you come to you come to you all the four is Minimax M3 is all the four is Minimax M3 is all the four is Minimax M3 is multimodal, M2.7, all the ones before multimodal, M2.7, all the ones before multimodal, M2.7, all the ones before that were not multimodal. And like, can that were not multimodal. And like, can that were not multimodal. And like, can you talk a little bit about the you talk a little bit about the you talk a little bit about the optimizations and optimizations and optimizations and how you trained it for that aspect? And how you trained it for that aspect? And how you trained it for that aspect? And then also, I want to get into the then also, I want to get into the then also, I want to get into the architecture of it afterwards as well.

  10. architecture of it afterwards as well. architecture of it afterwards as well. >> Right. Definitely. So, what's different >> Right. Definitely. So, what's different >> Right. Definitely. So, what's different from before was that it was trained from before was that it was trained from before was that it was trained multimodal from scratch. So, from step multimodal from scratch. So, from step multimodal from scratch. So, from step zero, we trained not only text data, we zero, we trained not only text data, we zero, we trained not only text data, we also trained image data. And it was also trained image data. And it was also trained image data. And it was normal for many other labs that the normal for many other labs that the normal for many other labs that the model would collapse after training a model would collapse after training a model would collapse after training a little bit, and we managed to solve that little bit, and we managed to solve that little bit, and we managed to solve that problem. And what we found was actually problem. And what we found was actually problem. And what we found was actually that with this kind of training from that with this kind of training from that with this kind of training from scratch, if you look at the attention scratch, if you look at the attention scratch, if you look at the attention map, it actually the visual of the the map, it actually the visual of the the map, it actually the visual of the the text tokens would attend to the visual text tokens would attend to the visual text tokens would attend to the visual tokens. So, that they are naturally tokens. So, that they are naturally tokens. So, that they are naturally combined together, they naturally combined together, they naturally combined together, they naturally understand each other. So, for example, understand each other. So, for example, understand each other. So, for example, we are developing, for example, we are developing, for example, we are developing, for example, websites, right? websites, right? websites, right? It is better if we train with It is better if we train with It is better if we train with both modalities. both modalities. both modalities. Also, like, for example, it can look at Also, like, for example, it can look at Also, like, for example, it can look at the website, it can understand how it the website, it can understand how it the website, it can understand how it looks, and then better optimize for it. looks, and then better optimize for it. looks, and then better optimize for it. Like for example, during reinforcement Like for example, during reinforcement Like for example, during reinforcement learning. learning. learning. Um so yeah, I think that is pretty cool. >> So >> So one one thing that kind of stuck out one one thing that kind of stuck out one one thing that kind of stuck out with this model for me was the fact that with this model for me was the fact that with this model for me was the fact that it introduced a lot of new things. The it introduced a lot of new things. The it introduced a lot of new things. The multi-modality, the increase of context multi-modality, the increase of context multi-modality, the increase of context to 1 million, um the fact that you have to 1 million, um the fact that you have to 1 million, um the fact that you have a sparse attention now. Um so if you go a sparse attention now. Um so if you go a sparse attention now. Um so if you go to the inference side, it's almost a a to the inference side, it's almost a a to the inference side, it's almost a a nightmare, isn't it? You get You get nightmare, isn't it? You get You get nightmare, isn't it? You get You get this new model and there's so many this new model and there's so many this new model and there's so many things that you could optimize to speed things that you could optimize to speed things that you could optimize to speed up inference. Um practically, what are up inference. Um practically, what are up inference. Um practically, what are the things that you focus on? There's the things that you focus on? There's the things that you focus on? There's like a thousand things that you could like a thousand things that you could like a thousand things that you could optimize, but where do you get the most optimize, but where do you get the most optimize, but where do you get the most bang for your buck?

  11. bang for your buck? bang for your buck? >> I mean, you focus on a thousand and one >> I mean, you focus on a thousand and one >> I mean, you focus on a thousand and one things, yeah. Like you just you just go things, yeah. Like you just you just go things, yeah. Like you just you just go and you keep you keep doing it. You find and you keep you keep doing it. You find and you keep you keep doing it. You find every edge that you can, um and and you every edge that you can, um and and you every edge that you can, um and and you go and and you push on it. Um so yeah, I go and and you push on it. Um so yeah, I go and and you push on it. Um so yeah, I think there there's there's no there's think there there's there's no there's think there there's there's no there's no stone that you leave unturned and you no stone that you leave unturned and you no stone that you leave unturned and you just just just uh keep keep going at it. If someone uh keep keep going at it. If someone uh keep keep going at it. If someone tells me you can't do the thousand first tells me you can't do the thousand first tells me you can't do the thousand first thing, yeah, like I don't know, try thing, yeah, like I don't know, try thing, yeah, like I don't know, try harder. harder. harder. >> Yeah. >> Yeah. >> Yeah. >> And then get up a little later. >> And then get up a little later. >> And then get up a little later. >> Yeah, like go ahead. >> Yeah, like go ahead. >> Yeah, like go ahead. >> No, I know. >> No, I know. >> No, I know. So the other thing that I wanted to ask So the other thing that I wanted to ask So the other thing that I wanted to ask is there's a whole kind of um is there's a whole kind of um is there's a whole kind of um zoo of open source models. As you're zoo of open source models. As you're zoo of open source models. As you're talking about speeding up inference, are talking about speeding up inference, are talking about speeding up inference, are there things that are there lessons that there things that are there lessons that there things that are there lessons that you can take from one model and apply it you can take from one model and apply it you can take from one model and apply it to MiniMax M3? Um or or do you have to to MiniMax M3? Um or or do you have to to MiniMax M3? Um or or do you have to like restart from scratch as you're like restart from scratch as you're like restart from scratch as you're thinking about the inference engine, the thinking about the inference engine, the thinking about the inference engine, the kernels, like how does that work? kernels, like how does that work? kernels, like how does that work? >> Right. Um yeah, so there there's >> Right. Um yeah, so there there's >> Right. Um yeah, so there there's definitely things that you learn from definitely things that you learn from definitely things that you learn from optimizing one model that you take to optimizing one model that you take to optimizing one model that you take to another. Um so I think sparse attentions another. Um so I think sparse attentions another. Um so I think sparse attentions are something that have become quite are something that have become quite are something that have become quite popular now. So the MiniMax sparse popular now. So the MiniMax sparse popular now. So the MiniMax sparse attention is a little bit different from attention is a little bit different from attention is a little bit different from the the the uh from the deep seek and the and and uh from the deep seek and the and and uh from the deep seek and the and and and those and the and the ones that you and those and the and the ones that you and those and the and the ones that you find in JLM. Um but the they're still find in JLM. Um but the they're still find in JLM. Um but the they're still similar lessons that you can take from similar lessons that you can take from similar lessons that you can take from that optimization process and that that optimization process and that that optimization process and that kernel writing process that you can then kernel writing process that you can then kernel writing process that you can then bring to to the new sparse attentions.

  12. bring to to the new sparse attentions. bring to to the new sparse attentions. And you know, we've been in in some form And you know, we've been in in some form And you know, we've been in in some form or or or or another I've been thinking about this or another I've been thinking about this or another I've been thinking about this problem for for many years. problem for for many years. problem for for many years. So going all the way back to my PhD. So So going all the way back to my PhD. So So going all the way back to my PhD. So it's it's it's great to see some it's it's it's great to see some it's it's it's great to see some validation that that folks can now train validation that that folks can now train validation that that folks can now train it at scale and people are using it and it at scale and people are using it and it at scale and people are using it and and it's and it's it's going pretty and it's and it's it's going pretty and it's and it's it's going pretty well. well. well. >> Yeah. >> Yeah. >> Yeah. >> Yeah, one of the interesting things >> Yeah, one of the interesting things >> Yeah, one of the interesting things especially about open source is you've especially about open source is you've especially about open source is you've got all these labs that are learning got all these labs that are learning got all these labs that are learning from each other kind of taking the wins from each other kind of taking the wins from each other kind of taking the wins from each other, right? So if if if one from each other, right? So if if if one from each other, right? So if if if one lab figures out that Minimax does this lab figures out that Minimax does this lab figures out that Minimax does this really well, then that gets becomes the really well, then that gets becomes the really well, then that gets becomes the golden standard. golden standard. golden standard. Um Um Um one of the things on model launch that I one of the things on model launch that I one of the things on model launch that I in the blog post that you guys go into in the blog post that you guys go into in the blog post that you guys go into is that this model was actually able to is that this model was actually able to is that this model was actually able to replicate a 12-hour run where it could replicate a 12-hour run where it could replicate a 12-hour run where it could reproduce an ICLR paper. reproduce an ICLR paper. reproduce an ICLR paper. And so somebody who's training this And so somebody who's training this And so somebody who's training this model to do this thing, how do you model to do this thing, how do you model to do this thing, how do you actually go about that? Cuz that seems actually go about that? Cuz that seems actually go about that? Cuz that seems like a pretty ludicrous task. like a pretty ludicrous task. like a pretty ludicrous task. >> Right. So letting the model to do cool >> Right. So letting the model to do cool >> Right. So letting the model to do cool stuff like replicating papers, stuff like replicating papers, stuff like replicating papers, optimizing kernel frameworks and stuff optimizing kernel frameworks and stuff optimizing kernel frameworks and stuff like that is always exciting for us like that is always exciting for us like that is always exciting for us researchers because it's like very researchers because it's like very researchers because it's like very related to our job. But training it can related to our job. But training it can related to our job. But training it can be very tricky because it's very long be very tricky because it's very long be very tricky because it's very long horizon and like horizon and like horizon and like the task itself would require GPUs.

  13. the task itself would require GPUs. the task itself would require GPUs. It has hardware constraints. It has hardware constraints. It has hardware constraints. So it's very interesting to train tasks So it's very interesting to train tasks So it's very interesting to train tasks like that and I would say the key there like that and I would say the key there like that and I would say the key there is still the environments and the data is still the environments and the data is still the environments and the data and how you formulate the problem, how and how you formulate the problem, how and how you formulate the problem, how you formulate the rewards, how you you formulate the rewards, how you you formulate the rewards, how you formulate the environment and how you formulate the environment and how you formulate the environment and how you change the reinforcement learning change the reinforcement learning change the reinforcement learning algorithm a little bit so that it's algorithm a little bit so that it's algorithm a little bit so that it's trained more efficiently trained more efficiently trained more efficiently so that you can see cool things emerging so that you can see cool things emerging so that you can see cool things emerging through the the iterations of RL runs. through the the iterations of RL runs. through the the iterations of RL runs. >> Mhm. >> Mhm. >> Mhm. Can you talk maybe if I keep pulling on Can you talk maybe if I keep pulling on Can you talk maybe if I keep pulling on this thread a little bit? How do you do this thread a little bit? How do you do this thread a little bit? How do you do evaluation over these longer longer evaluation over these longer longer evaluation over these longer longer scale runs? So if you're if you want the scale runs? So if you're if you want the scale runs? So if you're if you want the thing to do a 12-hour task, yes, it thing to do a 12-hour task, yes, it thing to do a 12-hour task, yes, it might or might not do it at the end, but might or might not do it at the end, but might or might not do it at the end, but are there like a intermediate things are there like a intermediate things are there like a intermediate things that you can also look at? that you can also look at? that you can also look at? >> Yes, we do. For these tasks, they are >> Yes, we do. For these tasks, they are >> Yes, we do. For these tasks, they are there are iterations, right? So, the there are iterations, right? So, the there are iterations, right? So, the model can submit several times, and we model can submit several times, and we model can submit several times, and we would evaluate each of them. Some of would evaluate each of them. Some of would evaluate each of them. Some of them some of the times the models would them some of the times the models would them some of the times the models would hack, and we do do like validation and hack, and we do do like validation and hack, and we do do like validation and test with for it to test if it's really test with for it to test if it's really test with for it to test if it's really improving on the performance or it's improving on the performance or it's improving on the performance or it's hacking.

  14. hacking. hacking. And also, we design our internal And also, we design our internal And also, we design our internal evaluations. So, for example, for the evaluations. So, for example, for the evaluations. So, for example, for the release of 2.7, release of 2.7, release of 2.7, we touched a bit on self-evolution, we touched a bit on self-evolution, we touched a bit on self-evolution, right? So, we're actively using the right? So, we're actively using the right? So, we're actively using the model to improve the speed of model to improve the speed of model to improve the speed of development internally, which development internally, which development internally, which like out of it out of it we can build like out of it out of it we can build like out of it out of it we can build our own evaluations that are closely our own evaluations that are closely our own evaluations that are closely related to our own work that we can related to our own work that we can related to our own work that we can evaluate the models on. evaluate the models on. evaluate the models on. >> Yeah. >> Yeah. >> Yeah. When So, when we start talking about When So, when we start talking about When So, when we start talking about these long horizon tasks that are 12 these long horizon tasks that are 12 these long horizon tasks that are 12 hours long, hours long, hours long, we gave an entire workshop on this on we gave an entire workshop on this on we gave an entire workshop on this on Monday, but what I wanted to come to Monday, but what I wanted to come to Monday, but what I wanted to come to you, Dan, for is you, Dan, for is you, Dan, for is KB cache. So, if let's say you you have KB cache. So, if let's say you you have KB cache. So, if let's say you you have concurrent requests that are 500 to a concurrent requests that are 500 to a concurrent requests that are 500 to a million million million thousand context length long, how do you thousand context length long, how do you thousand context length long, how do you deal with the KB cache that just keeps deal with the KB cache that just keeps deal with the KB cache that just keeps on growing, and how does the on growing, and how does the on growing, and how does the infrastructure deal with that? infrastructure deal with that? infrastructure deal with that? >> Yeah, so there there's a lot of >> Yeah, so there there's a lot of >> Yeah, so there there's a lot of different pieces that that you put different pieces that that you put different pieces that that you put there. Like in in some sense, it's like there. Like in in some sense, it's like there. Like in in some sense, it's like recreating a distributed file system. recreating a distributed file system. recreating a distributed file system. So, we we're in some sense building So, we we're in some sense building So, we we're in some sense building something like that something like that something like that or a very big database. or a very big database. or a very big database. It's it's pretty simple in theory. It's It's it's pretty simple in theory. It's It's it's pretty simple in theory. It's like the type of thing that you should like the type of thing that you should like the type of thing that you should have done in your third year of have done in your third year of have done in your third year of undergrad or something like that, but undergrad or something like that, but undergrad or something like that, but most of us actually skipped that class, most of us actually skipped that class, most of us actually skipped that class, so now we're rediscovering it live in in so now we're rediscovering it live in in so now we're rediscovering it live in in industry. But it's it's all about where industry. But it's it's all about where industry. But it's it's all about where do you store that cache?

  15. do you store that cache? do you store that cache? How do you know Have you have you seen How do you know Have you have you seen How do you know Have you have you seen this before? How do you fetch it? How do this before? How do you fetch it? How do this before? How do you fetch it? How do how do you send it from one place to how do you send it from one place to how do you send it from one place to another? another? another? So, yeah, it's it's a it's it's not that So, yeah, it's it's a it's it's not that So, yeah, it's it's a it's it's not that complicated, but you you do have to make complicated, but you you do have to make complicated, but you you do have to make sure that that you do a good job. sure that that you do a good job. sure that that you do a good job. >> Yeah, one thing I noticed you you gave a >> Yeah, one thing I noticed you you gave a >> Yeah, one thing I noticed you you gave a a um a lecture at Stanford recently and a um a lecture at Stanford recently and a um a lecture at Stanford recently and one thing that stood out was that if you one thing that stood out was that if you one thing that stood out was that if you fast forward like 2 3 years, 2 3 years fast forward like 2 3 years, 2 3 years fast forward like 2 3 years, 2 3 years is is is a long time in AI. And if you is is is a long time in AI. And if you is is is a long time in AI. And if you look back you said that we'll realize look back you said that we'll realize look back you said that we'll realize that how um how early we are right now. that how um how early we are right now. that how um how early we are right now. >> Yeah. >> Yeah. >> Yeah. >> So, from your vantage point, where 3 >> So, from your vantage point, where 3 >> So, from your vantage point, where 3 years out, what do you think we'll look years out, what do you think we'll look years out, what do you think we'll look back on be like, "Why were we doing it back on be like, "Why were we doing it back on be like, "Why were we doing it this way?" this way?" this way?" >> Great question. Um uh some things that I >> Great question. Um uh some things that I >> Great question. Um uh some things that I hope for, so I think we underutilize our hope for, so I think we underutilize our hope for, so I think we underutilize our GPUs a lot right now. Um you know, GPUs a lot right now. Um you know, GPUs a lot right now. Um you know, SpaceX said they would have like 10% SpaceX said they would have like 10% SpaceX said they would have like 10% flop utilization or something like that. flop utilization or something like that. flop utilization or something like that. I hope in 3 years, well, they they I hope in 3 years, well, they they I hope in 3 years, well, they they should already be embarrassed about it, should already be embarrassed about it, should already be embarrassed about it, but I hope in 3 years they're extra but I hope in 3 years they're extra but I hope in 3 years they're extra embarrassed by it. Um so, certainly embarrassed by it. Um so, certainly embarrassed by it. Um so, certainly training should be should be pretty training should be should be pretty training should be should be pretty um should be pretty good. I think at um should be pretty good. I think at um should be pretty good. I think at Anthropic we can do a lot better with Anthropic we can do a lot better with Anthropic we can do a lot better with the hardware that we're using uh that we the hardware that we're using uh that we the hardware that we're using uh that we have today, so I hope uh in a few years have today, so I hope uh in a few years have today, so I hope uh in a few years we'll we'll say uh we'll we'll have seen we'll we'll say uh we'll we'll have seen we'll we'll say uh we'll we'll have seen the light on on some of those pieces. Um the light on on some of those pieces. Um the light on on some of those pieces. Um and I think there'll be a lot more and I think there'll be a lot more and I think there'll be a lot more models, there'll be a lot better. Um I I models, there'll be a lot better. Um I I models, there'll be a lot better. Um I I hope finally by then we've put uh hope finally by then we've put uh hope finally by then we've put uh put to bed this question about the open put to bed this question about the open put to bed this question about the open models. You know, there's every few models. You know, there's every few models. You know, there's every few months there's someone like, "Oh, months there's someone like, "Oh, months there's someone like, "Oh, Anthropic OpenAI, they're so ahead, yada Anthropic OpenAI, they're so ahead, yada Anthropic OpenAI, they're so ahead, yada yada." Um but I think uh we're we're yada." Um but I think uh we're we're yada." Um but I think uh we're we're seeing with models like M3 and GLM and seeing with models like M3 and GLM and seeing with models like M3 and GLM and Kimmy and and all those models that um

  16. Kimmy and and all those models that um Kimmy and and all those models that um the open-source frontier really can the open-source frontier really can the open-source frontier really can catch up. Um and and it's it's not even catch up. Um and and it's it's not even catch up. Um and and it's it's not even that far behind, so so I think that that far behind, so so I think that that far behind, so so I think that that's quite exciting. that's quite exciting. that's quite exciting. >> Yeah. >> Yeah. >> Yeah. I I wanted to throw that same question I I wanted to throw that same question I I wanted to throw that same question over to you, Olive, but you mentioned over to you, Olive, but you mentioned over to you, Olive, but you mentioned that you it for for the M2 series and that you it for for the M2 series and that you it for for the M2 series and the M3 series, you're using this idea of the M3 series, you're using this idea of the M3 series, you're using this idea of self-evolution, where the model is self-evolution, where the model is self-evolution, where the model is building its own harness, and then it's building its own harness, and then it's building its own harness, and then it's training inside of that, and then you training inside of that, and then you training inside of that, and then you get the next checkpoint. Um get the next checkpoint. Um get the next checkpoint. Um if you look back 3 years uh out and then if you look back 3 years uh out and then if you look back 3 years uh out and then you say like, "What in RL or you say like, "What in RL or you say like, "What in RL or post-training um do you think uh made post-training um do you think uh made post-training um do you think uh made the biggest difference?" What do you the biggest difference?" What do you the biggest difference?" What do you think that is from this vantage point? think that is from this vantage point? think that is from this vantage point? >> Um great question. But, 3 years ago I >> Um great question. But, 3 years ago I >> Um great question. But, 3 years ago I was still in in was still in in was still in in >> [laughter] >> [laughter] >> [laughter] >> I actually didn't start this industry >> I actually didn't start this industry >> I actually didn't start this industry yet. yet. yet. >> Yeah. So I wouldn't have imagined what's >> Yeah. So I wouldn't have imagined what's >> Yeah. So I wouldn't have imagined what's happening right now today. So it's happening right now today. So it's happening right now today. So it's really exciting. But I can see how really exciting. But I can see how really exciting. But I can see how models that were developed are were models that were developed are were models that were developed are were already improving the speed of already improving the speed of already improving the speed of development maybe a year ago or even development maybe a year ago or even development maybe a year ago or even further than a year ago. So I could see further than a year ago. So I could see further than a year ago. So I could see how this speed is actually accelerating. how this speed is actually accelerating. how this speed is actually accelerating. How the development is accelerating. And How the development is accelerating. And How the development is accelerating. And that's how like open weight models can that's how like open weight models can that's how like open weight models can really catch up with frontier labs and really catch up with frontier labs and really catch up with frontier labs and Yeah, Yeah, Yeah, that's how we think we are more mission that's how we think we are more mission that's how we think we are more mission to bring this model to everyone to bring this model to everyone to bring this model to everyone so that everyone can use it. Yeah.

  17. so that everyone can use it. Yeah. so that everyone can use it. Yeah. >> Awesome. Thank you guys. Thank you Dan. >> Awesome. Thank you guys. Thank you Dan. >> Awesome. Thank you guys. Thank you Dan. Thank you Thank you Olive. Thank you Thank you Thank you Olive. Thank you Thank you Thank you Olive. Thank you guys so much. Have a great day. Very guys so much. Have a great day. Very guys so much. Have a great day. Very cool. Thanks so much.

Summary

This discussion explores the rapid advancement of AI, focusing on the open-sourcing of MiniMax's M3 model. Key subjects include GPU optimization for AI inference and the benefits of community contribution to model development. The practical takeaway is that open-sourcing powerful AI models fosters collective improvement and broader accessibility, enabling faster inference and better serving of AI applications.

View original episode ↗