Why AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, MiniMax
Read full transcript 18 segments
-
Joining us on stage is the Joining us on stage is the co-founder and co-founder and co-founder and chief scientific chief scientific chief scientific officer of Hugging Face, officer of Hugging Face, officer of Hugging Face, Thomas Wolfe. Thomas Wolfe. Greetings to everyone. Congratulations, Greetings to everyone. Congratulations, Olive. Glad to see Olive. Glad to see Olive. Glad to see you on stage. you on stage. you on stage. Hello, bye. Hello, bye. Hello, bye. Thank you for inviting me. Thank you for inviting me. Thank you for inviting me. I think you're in for I think you're in for I think you're in for something something something interesting today, because you just interesting today, because you just interesting today, because you just saw GLM, which is currently saw GLM, which is currently saw GLM, which is currently ranked second in the ranked second in the ranked second in the artificial artificial artificial intelligence rankings. I removed intelligence rankings. I removed intelligence rankings. I removed Fable because no one can Fable because no one can Fable because no one can use it. And use it. And use it. And now we have number now we have number now we have number four. So you four. So you four. So you will have all the best will have all the best will have all the best models, at least the models, at least the models, at least the best best best open source models. And open source models. And open source models. And we are very lucky to have we are very lucky to have we are very lucky to have Olive with us, who Olive with us, who Olive with us, who has had a pretty has had a pretty has had a pretty amazing life amazing life amazing life path. She came path. She came path. She came to the USA, to to the USA, to to the USA, to Pennsylvania. She was a graduate student at Pennsylvania. She was a graduate student at NYU at the NYU at the Yandex lab, Yandex lab, Yandex lab, working on J-pop, but working on J-pop, but working on J-pop, but we decided not to we decided not to we decided not to talk talk talk about J-pop today, right? This is for about J-pop today, right? This is for about J-pop today, right? This is for another time. And another time. And another time. And instead of instead of instead of joining Hugging joining Hugging joining Hugging Face, which Face, which Face, which was also in New York at the time, she was also in New York at the time, she was also in New York at the time, she decided to join decided to join decided to join MiniMax. For those of you who MiniMax. For those of you who MiniMax. For those of you who may not know all the may not know all the may not know all the AI labs AI labs AI labs around the around the around the world, you're world, you're world, you're forgiven, because there are probably about 64 of them right now.
-
forgiven, because there are probably about 64 of them right now. MiniMax is one of the MiniMax is one of the leading companies leading companies leading companies that we call "AI that we call "AI dragons" in China. These are the dragons" in China. These are the dragons" in China. These are the new players: Deep Seek, new players: Deep Seek, new players: Deep Seek, which is now very well known which is now very well known , Moonshot, which created Kimi, , Moonshot, which created Kimi, , Moonshot, which created Kimi, Z and GLM, which you just Z and GLM, which you just Z and GLM, which you just saw, and now we saw, and now we saw, and now we have MiniMax. They are all have MiniMax. They are all have MiniMax. They are all extremely good, extremely good, extremely good, these are extremely these are extremely these are extremely talented teams talented teams talented teams fighting for first fighting for first fighting for first place. The latest place. The latest place. The latest MiniMax release is the M3, MiniMax release is the M3, MiniMax release is the M3, introduced in introduced in introduced in early June, which early June, which early June, which was the was the was the best best best open source model at the time. Very open source model at the time. Very open source model at the time. Very impressive. There are impressive. There are impressive. There are a lot of a lot of a lot of interesting things interesting things interesting things to tell about these models, so we'll to tell about these models, so we'll to tell about these models, so we'll quickly dive into the quickly dive into the quickly dive into the details. And then details. And then details. And then let's talk a little about let's talk a little about let's talk a little about what exactly makes what exactly makes what exactly makes MiniMax special, what's MiniMax special, what's MiniMax special, what's interesting about it. interesting about it. Maybe I'll give Maybe I'll give Olive the floor so she can Olive the floor so she can Olive the floor so she can start a little. Can you, you start a little. Can you, you start a little. Can you, you know, share know, share know, share your vision for the M3, what your vision for the M3, what your vision for the M3, what you like about this you like about this you like about this model, how the model, how the model, how the release went? release went? release went? Mhm. Yes, we Mhm. Yes, we Mhm. Yes, we released M3 earlier released M3 earlier released M3 earlier this month, it's a this month, it's a this month, it's a smaller model with a smaller model with a smaller model with a total number of total number of total number of parameters of about 400 parameters of about 400 parameters of about 400 billion and 20 billion and 20 billion and 20 billion billion billion activated. But it is activated. But it is activated. But it is very effective both in terms of very effective both in terms of very effective both in terms of encoding and encoding and encoding and understanding understanding understanding visual data. So what visual data. So what open source models typically don't have is that open source models typically don't have is that this model can this model can this model can not only work with not only work with not only work with code, but it can understand code, but it can understand code, but it can understand video, images, and video, images, and video, images, and it has a super long it has a super long it has a super long context of 1 million.
-
context of 1 million. Um, with our new Um, with our new architecture called architecture called architecture called MSA, Minimax Sparse MSA, Minimax Sparse MSA, Minimax Sparse Attention. We really Attention. We really Attention. We really combined these three things combined these three things combined these three things because we know they're going to be because we know they're going to be because we know they're going to be really really really important for important for important for future AI future AI applications: applications: applications: coding capabilities coding capabilities , agent capabilities, , agent capabilities, , agent capabilities, longer context, and longer context, and longer context, and multimodal multimodal multimodal understanding. Um, yes, understanding. Um, yes, understanding. Um, yes, I think that would be very I think that would be very I think that would be very interesting regarding this interesting regarding this interesting regarding this model. model. model. Yes, there's a lot to Yes, there's a lot to Yes, there's a lot to take in this take in this take in this model, and it seems to be model, and it seems to be the only model out of the top 5 the only model out of the top 5 the only model out of the top 5 in the public domain in the public domain in the public domain that is actually that is actually that is actually multimodal, multimodal, multimodal, so it's worth so it's worth so it's worth talking about. But talking about. But talking about. But maybe first about the maybe first about the maybe first about the long context, because long context, because long context, because this was the first model this was the first model this was the first model that really had this that really had this that really had this real context for real context for real context for 1 million tokens that 1 million tokens that 1 million tokens that really works. And really works. And really works. And you also had Minimax you also had Minimax you also had Minimax Sparse Attention—that's a Sparse Attention—that's a Sparse Attention—that's a technique for technique for technique for improving improving improving performance that you performance that you performance that you also published also published also published and widely and widely and widely distributed. So distributed. So distributed. So can you can you can you talk a little bit about that?
-
talk a little bit about that? talk a little bit about that? Perhaps about how Perhaps about how Perhaps about how the project evolved the project evolved the project evolved from an attention mechanism from an attention mechanism from an attention mechanism to creating such a to creating such a to creating such a long context. long context. long context. Yes. I would say that Yes. I would say that Yes. I would say that the history of long the history of long the history of long context goes back to context goes back to context goes back to the days of Minimax M1 and Minimax 01, the days of Minimax M1 and Minimax 01, the days of Minimax M1 and Minimax 01, where the model where the model where the model was actually capable of was actually capable of was actually capable of performing tasks performing tasks performing tasks with a context of 10 with a context of 10 with a context of 10 million tokens. million tokens. million tokens. 10 million? 10 million? 10 million? 10 million, yes. But 10 million, yes. But 10 million, yes. But then it wasn't an then it wasn't an then it wasn't an agent model, right? agent model, right? agent model, right? It was just, It was just, It was just, for example, for example, for example, downloading a book and downloading a book and downloading a book and she could give she could give she could give reviews on it and so on. reviews on it and so on. So we realized that So we realized that longer context longer context longer context actually opens up actually opens up actually opens up a lot of possibilities, a lot of possibilities, a lot of possibilities, especially when especially when especially when interacting with interacting with interacting with users. And users. And users. And now that the agent now that the agent now that the agent interacts with the entire interacts with the entire interacts with the entire environment, receives environment, receives environment, receives all the responses from the all the responses from the all the responses from the tools, and has tools, and has tools, and has multi-round multi-round multi-round dialogues, a shorter dialogues, a shorter dialogues, a shorter context would not be context would not be context would not be enough to enough to enough to complete complex complete complex complete complex tasks. So for this tasks. So for this tasks. So for this version we said, "Oh, version we said, "Oh, version we said, "Oh, we should have we should have we should have support for longer support for longer support for longer context." So, we context." So, we context." So, we went the route of went the route of went the route of using our using our using our Minimax Sparse Attention. This is Minimax Sparse Attention. This is Minimax Sparse Attention. This is an architecture that was an architecture that was an architecture that was scalable and had a scalable and had a scalable and had a simple design.
-
simple design. In general, at a high In general, at a high level, it has an level, it has an level, it has an index branch that, index branch that, index branch that, at a general level, at a general level, at a general level, selects what is more selects what is more selects what is more important in the important in the important in the context. And then context. And then context. And then we have a we have a sparse attention branch, which sparse attention branch, which performs computations performs computations performs computations on selected blocks on selected blocks on selected blocks to directly to directly to directly perform the task. perform the task. So yes, we really So yes, we really designed an elegant designed an elegant designed an elegant architecture so that we architecture so that we architecture so that we could scale could scale could scale the length, and in the length, and in the length, and in the future, the the future, the the future, the size of the model. size of the model. This is great. I This is great. I like how, for like how, for like how, for those who have been in this industry for a long time those who have been in this industry for a long time those who have been in this industry for a long time , we've worked a lot , we've worked a lot , we've worked a lot on the on the on the attention mechanism, right? attention mechanism, right? attention mechanism, right? This is n-square, and there were This is n-square, and there were This is n-square, and there were many variations of many variations of many variations of linear attention. linear attention. linear attention. Yes. Yes. Yes. And then somehow all of that And then somehow all of that And then somehow all of that disappeared when disappeared when disappeared when Flash Attention appeared. We found Flash Attention appeared. We found Flash Attention appeared. We found that we simply that we simply that we simply needed a more needed a more needed a more efficient core. Now efficient core. Now efficient core. Now I like how I like how I like how we're going back to we're going back to we're going back to thinking, you know, about the thinking, you know, about the thinking, you know, about the basics: what is basics: what is basics: what is attention? How can we attention? How can we attention? How can we make this make this make this more efficient? more efficient? more efficient? A million tokens is A million tokens is A million tokens is crazy, right? GPT-2 crazy, right? GPT-2 crazy, right? GPT-2 had 1024, and everyone said, had 1024, and everyone said, had 1024, and everyone said, "Oh, that's a lot.
-
"Oh, that's a lot. "Oh, that's a lot. We'll never We'll never We'll never need more." need more." need more." How do you see the future of How do you see the future of How do you see the future of this direction? Jeff this direction? Jeff this direction? Jeff Dean offered Dean offered Dean offered me attention for a me attention for a me attention for a trillion tokens the other day. Do you trillion tokens the other day. Do you trillion tokens the other day. Do you think we should think we should think we should go to a trillion? go to a trillion? This is definitely something we This is definitely something we can explore can explore can explore in the direction of an in the direction of an in the direction of an ultra-long ultra-long ultra-long context. This is a very context. This is a very context. This is a very exciting area for exciting area for exciting area for research that research that research that will require will require will require a lot of work on the a lot of work on the a lot of work on the design of the architecture design of the architecture design of the architecture in combination with in combination with in combination with the hardware. Yes. Do the hardware. Yes. Do the hardware. Yes. Do you think there is still you think there is still you think there is still much that could be much that could be much that could be easily improved? We easily improved? We easily improved? We saw OpenAI significantly saw OpenAI significantly saw OpenAI significantly reduce—we don't reduce—we don't reduce—we don't know how, but they know how, but they know how, but they reduced—the reduced—the reduced—the inference costs by half, inference costs by half, inference costs by half, probably due to more probably due to more probably due to more efficient efficient efficient attention processing. Do you think there are still attention processing. Do you think there are still attention processing. Do you think there are still many simple many simple many simple solutions to how we solutions to how we solutions to how we handle this? One handle this? One handle this? One interesting thing about the M3 is that interesting thing about the M3 is that interesting thing about the M3 is that it's cheap, it's cheap, it's cheap, partly because of its partly because of its partly because of its sparse attention or sparse attention or , partly, because of its , partly, because of its , partly, because of its compactness, but it's compactness, but it's compactness, but it's also very also very also very efficient. efficient. efficient. Yes. Yes. Yes. Do you think we Do you think we Do you think we can go even further? And can go even further? And maybe how did you maybe how did you maybe how did you invent MinMax fast attention?
-
invent MinMax fast attention? invent MinMax fast attention? Was this an idea Was this an idea Was this an idea suggested by the suggested by the suggested by the agent? Was it still a agent? Was it still a agent? Was it still a human idea? human idea? human idea? Tell us a little Tell us a little Tell us a little about it. about it. about it. Yes. We believe that there is Yes. We believe that there is Yes. We believe that there is still a lot of work to be still a lot of work to be still a lot of work to be done in the done in the done in the architecture and architecture and architecture and optimization of the optimization of the optimization of the output to make the model output to make the model output to make the model more efficient, more efficient, more efficient, especially for tasks especially for tasks especially for tasks where high where high where high capabilities are required, right? For capabilities are required, right? For capabilities are required, right? For such tasks, we such tasks, we such tasks, we really want really want really want the model to be the model to be the model to be efficient. And who efficient. And who efficient. And who invented sparse invented sparse invented sparse attention? Actually, attention? Actually, attention? Actually, I think an I think an I think an intern from intern from intern from our team worked on this. Yes, an our team worked on this. Yes, an our team worked on this. Yes, an intern. This is usually not possible in intern. This is usually not possible in intern. This is usually not possible in many many many laboratories laboratories laboratories because because because interns there do not have interns there do not have interns there do not have access to data, access to data, access to data, work, etc. But yes, work, etc. But yes, work, etc. But yes, we are open to we are open to we are open to anyone who wants to anyone who wants to anyone who wants to contribute to contribute to contribute to our models. So, our models. So, our models. So, the architecture the architecture the architecture was actually designed by an was actually designed by an was actually designed by an intern. intern. intern. This is very good. There This is very good. There This is very good. There is still work here for interns is still work here for interns is still work here for interns . Good, good . Good, good . Good, good news. It's also a news. It's also a news. It's also a great transition to how great transition to how great transition to how Min Max works Min Max works Min Max works from the inside. We from the inside. We from the inside. We discussed this discussed this discussed this before going on before going on before going on stage, I said that stage, I said that stage, I said that anyone can anyone can anyone can propose a project propose a project . Can you . Can you . Can you tell me a little bit about tell me a little bit about tell me a little bit about how you are how you are how you are organized, how you organized, how you organized, how you conduct your conduct your conduct your research?
-
research? research? Mhm. Mhm. I think it's Mhm. Mhm. I think it's Mhm. Mhm. I think it's very different very different very different from school from school from school or even early or even early or even early tech tech tech companies. This is quite, companies. This is quite, companies. This is quite, quite excellent. We quite excellent. We quite excellent. We make sure to make sure to make sure to have a good foundation and have a good foundation and have a good foundation and infrastructure so that infrastructure so that infrastructure so that everyone can everyone can everyone can experiment with the experiment with the experiment with the model and think about model and think about model and think about what can be improved in it what can be improved in it what can be improved in it . And then, . And then, . And then, after the model is released, after the model is released, after the model is released, when they are free, when they are free, when they are free, right? They can right? They can right? They can work with the model. work with the model. work with the model. They can They can They can come up with their own come up with their own come up with their own assessment methods. They assessment methods. They assessment methods. They can find weaknesses can find weaknesses can find weaknesses and and and suggest things they suggest things they suggest things they want to improve in the want to improve in the want to improve in the model. Then other model. Then other model. Then other people who are interested people who are interested people who are interested in it can in it can in it can join the join the join the project and they project and they project and they will work on will work on will work on it for a few weeks or it for a few weeks or it for a few weeks or even a few even a few even a few months. When everything months. When everything months. When everything works out, the final works out, the final works out, the final result result result is implemented into is implemented into is implemented into our model. We our model. We our model. We use this in use this in use this in our final our final our final training, and it becomes training, and it becomes training, and it becomes available to the available to the available to the audience. audience. audience. That is, people can That is, people can That is, people can work on a work on a work on a project for a really project for a really project for a really long time. When you say long time. When you say long time. When you say "a few months," that "a few months," that "a few months," that can be a really can be a really can be a really deep exploration of deep exploration of deep exploration of what's possible.
-
what's possible. what's possible. Yes. I would say, Yes. I would say, Yes. I would say, for example, for example, for example, architecture may architecture may architecture may require longer require longer require longer time for research, time for research, time for research, analysis, experimentation, analysis, experimentation, analysis, experimentation, even re- even re- even re- conducting conducting conducting assessments for assessments for assessments for prior prior prior learning. Yes, so it learning. Yes, so it learning. Yes, so it may take more may take more may take more time. time. time. It's very nice, yes. It's very nice, yes. It's very nice, yes. And I know you also And I know you also And I know you also pay a lot of pay a lot of pay a lot of attention to evaluation. I attention to evaluation. I attention to evaluation. I agree. We can agree. We can agree. We can talk about it. One talk about it. One talk about it. One thing that is probably thing that is probably thing that is probably related to this is the related to this is the related to this is the unique unique unique feature of M3 and feature of M3 and feature of M3 and your team regarding your team regarding your team regarding multimodality. multimodality. multimodality. So not only text, but the So not only text, but the So not only text, but the same model can also same model can also same model can also understand images understand images understand images and videos. As far as I and videos. As far as I and videos. As far as I understand, but understand, but understand, but please explain better please explain better . When you read the . When you read the . When you read the model card on Hugging Face, it model card on Hugging Face, it model card on Hugging Face, it says that the model says that the model says that the model was trained from the first was trained from the first was trained from the first step as step as step as multimodal, not multimodal, not multimodal, not just as an added-on just as an added-on just as an added-on later, right? later, right? later, right? Can you talk Can you talk Can you talk a little more about this, a little more about this, a little more about this, why you think it's why you think it's why you think it's important, and why important, and why important, and why starting with the first starting with the first starting with the first step on step on step on multimodal multimodal multimodal learning is better than learning is better than learning is better than just training a just training a just training a model. model. model. So, we call it our So, we call it our So, we call it our own own own multimodality. multimodality. Typically, model development labs train model development labs train multimodal multimodal multimodal capabilities, such as capabilities, such as visual pattern recognition, visual pattern recognition, after completing after completing after completing prior prior prior training with text.
-
training with text. They add They add adapters and then adapters and then adapters and then train that part. train that part. train that part. But we found that But we found that But we found that this actually hurts this actually hurts this actually hurts the performance of the performance of the performance of text models. In text models. In text models. In addition, the performance of addition, the performance of visual data recognition does not visual data recognition does not stabilize stabilize stabilize properly properly properly because the model because the model because the model focuses only on focuses only on focuses only on text understanding text understanding , which is not an , which is not an , which is not an optimal solution optimal solution . And, if you think about it, . And, if you think about it, . And, if you think about it, it's not very it's not very it's not very scalable either. We scalable either. We scalable either. We want to scale want to scale want to scale data, right? Some data, right? Some data, right? Some labs also labs also labs also train this train this train this ability starting ability starting ability starting midway through the midway through the midway through the training process. For example, training process. For example, training process. For example, by continuing by continuing by continuing previous training. previous training. previous training. But we found that it But we found that it But we found that it depends a lot on depends a lot on depends a lot on the approach, the the approach, the the approach, the “recipe,” so to speak. “recipe,” so to speak. For different For different architectures, different architectures, different architectures, different datasets, and datasets, and datasets, and different different different learning rates, the recipes learning rates, the recipes learning rates, the recipes will be different. will be different. It is difficult It is difficult to control, it is difficult to control, it is difficult to control, it is difficult to scale up to scale up to scale up the results of the results of the results of experiments, and it experiments, and it experiments, and it is impossible to transfer is impossible to transfer is impossible to transfer the findings to a larger the findings to a larger the findings to a larger model. So we model. So we model. So we thought: why not thought: why not thought: why not start learning from the start learning from the start learning from the very first step?
-
very first step? This seems the most This seems the most natural. We know natural. We know natural. We know that many that many that many labs labs labs face face face challenges doing challenges doing challenges doing this. The model this. The model this. The model broke down after broke down after broke down after a few training steps a few training steps trying to combine trying to combine text and visual text and visual text and visual understanding, but we understanding, but we understanding, but we managed to solve this managed to solve this managed to solve this problem. We have done problem. We have done problem. We have done a lot of work on VIT and a lot of work on VIT and a lot of work on VIT and on the data we on the data we on the data we train on. train on. For example, we For example, we use what we use what we use what we call call call data interleaving. This is data interleaving. This is essentially natural essentially natural essentially natural data: we leave the data: we leave the data: we leave the images and videos alone, images and videos alone, images and videos alone, instead of instead of instead of masking them, perform masking them, perform masking them, perform high-quality data cleaning and high-quality data cleaning and high-quality data cleaning and masking, masking, masking, and apply and apply and apply efficient efficient reward modeling to reward modeling to train the model from the train the model from the train the model from the first step—and it first step—and it scales very well. It scales very well. It does not collapse. does not collapse. does not collapse. This is truly impressive. This is truly impressive. This is truly impressive. Impressive. Should Impressive. Should Impressive. Should we expect we expect we expect significantly larger models significantly larger models significantly larger models in the future? This in the future? This in the future? This model is still pretty model is still pretty model is still pretty small, right? It small, right? It small, right? It has 428 billion has 428 billion has 428 billion parameters, of which 23 parameters, of which 23 parameters, of which 23 billion are active. Do billion are active. Do billion are active. Do you think you think you think you will surpass the you will surpass the you will surpass the trillion mark?
-
trillion mark? trillion mark? Certainly. Yes, Certainly. Yes, Certainly. Yes, definitely in definitely in definitely in the future. There are many the future. There are many the future. There are many tasks that a tasks that a tasks that a model will not be able to model will not be able to model will not be able to handle handle handle properly with fewer properly with fewer properly with fewer parameters parameters . We are definitely planning . We are definitely planning . We are definitely planning much more ambitious much more ambitious much more ambitious projects. projects. projects. This is great. This is great. This is great. We are looking forward to it. We are looking forward to it. We are looking forward to it. Another interesting thing Another interesting thing Another interesting thing that always that always that always fascinates me about MiniMax is fascinates me about MiniMax is fascinates me about MiniMax is how you have developed a how you have developed a how you have developed a whole network of applications whole network of applications whole network of applications and products, and products, and products, right? So, I remember right? So, I remember right? So, I remember that MiniMax started that MiniMax started that MiniMax started publishing open publishing open publishing open source on the Hugging source on the Hugging source on the Hugging Face platform back in January of Face platform back in January of Face platform back in January of last year. That is, last year. That is, last year. That is, it was 18 months it was 18 months it was 18 months ago. We talked a little bit ago. We talked a little bit ago. We talked a little bit about about about your team to your team to your team to get a sense of what you get a sense of what you get a sense of what you do, and I do, and I do, and I remember you saying remember you saying remember you saying that you already have a that you already have a that you already have a huge number of huge number of huge number of users on users on users on some of these some of these some of these apps. Can you apps. Can you apps. Can you tell us a little about tell us a little about tell us a little about how it all started how it all started ? Was it the case that you ? Was it the case that you ? Was it the case that you already had a lot of already had a lot of already had a lot of applications and then applications and then applications and then you thought: we have you thought: we have you thought: we have all this data, why not all this data, why not all this data, why not train a model and train a model and train a model and then create a then create a then create a research team research team ? How did this ? How did this ? How did this story develop? story develop? story develop? Our story Our story has been built around a has been built around a model from day one. I believe that model from day one. I believe that model from day one. I believe that multimodality—a multimodality—a model that can model that can model that can understand all understand all understand all visual images and visual images and visual images and deliver results deliver results deliver results in any in any in any modality—was the modality—was the modality—was the first goal our first goal our first goal our CEO planned before the CEO planned before the CEO planned before the company was founded.
-
company was founded. company was founded. So it was a dream of So it was a dream of So it was a dream of general-purpose AI. general-purpose AI. general-purpose AI. I think it was very I think it was very I think it was very early, even before early, even before early, even before ChatGPT appeared. ChatGPT appeared. ChatGPT appeared. Wow. Wow. Wow. Yes. And apps Yes. And apps Yes. And apps came about as an application came about as an application came about as an application because when you have because when you have because when you have the capabilities of a model, you the capabilities of a model, you the capabilities of a model, you want people to be able to want people to be able to use them. Not use them. Not many people can many people can many people can use APIs, use APIs, use APIs, right? We can't right? We can't right? We can't expect everyone expect everyone expect everyone to be able to work with the API. to be able to work with the API. Therefore, we need Therefore, we need good interfaces, good interfaces, good interfaces, good applications, and good applications, and good applications, and interesting scenarios that interesting scenarios that allow people to experience allow people to experience the capabilities of the model. the capabilities of the model. These applications are believed to These applications are believed to have reached have reached have reached over 300 million over 300 million over 300 million users in users in users in about 200 countries around about 200 countries around about 200 countries around the world. And, I think, over a the world. And, I think, over a the world. And, I think, over a million companies million companies million companies too. too. too. Yes, it was Yes, it was Yes, it was amazing when I amazing when I amazing when I heard about such a heard about such a heard about such a scale. We often do not scale. We often do not scale. We often do not realize the real realize the real realize the real extent of such extent of such extent of such use. This use. This use. This brings me to brings me to brings me to the question of the the question of the open open open source business model—the eternal question: it's source business model—the eternal question: it's source business model—the eternal question: it's nice to make nice to make nice to make models open now, models open now, models open now, but you also but you also but you also need some need some need some sources of revenue, sources of revenue, sources of revenue, right? For example, you right? For example, you right? For example, you decided to provide N3 decided to provide N3 decided to provide N3 for free, and I for free, and I for free, and I think that's great think that's great think that's great for the world. How do you for the world. How do you for the world. How do you look at this? Do look at this? Do look at this? Do you have separate models you have separate models you have separate models that you that you that you use only use only use only for the application? Do for the application? Do for the application? Do you plan to you plan to continue to continue to make the models make the models make the models publicly available in the future? It's publicly available in the future? It's publicly available in the future? It's probably hard probably hard probably hard to say for sure, but are to say for sure, but are to say for sure, but are you planning on doing that? What is you planning on doing that? What is you planning on doing that? What is the culture the culture the culture around open
-
around open around open source now? source now? source now? Personally, I, like Personally, I, like Personally, I, like our entire team of our entire team of our entire team of researchers, always researchers, always researchers, always hope to make hope to make hope to make models open. This is models open. This is models open. This is our plan because we our plan because we our plan because we see how the see how the open source community open source community is helping is helping is helping to improve the model. to improve the model. For example, we For example, we get a lot of get a lot of get a lot of feedback on feedback on model performance from our model performance from our wonderful community, and wonderful community, and wonderful community, and we get PR for we get PR for we get PR for everything we everything we everything we discover, right? discover, right? discover, right? And they are very, very And they are very, very And they are very, very valuable, so we valuable, so we valuable, so we are taking them into account in are taking them into account in are taking them into account in future versions. future versions. future versions. So open source So open source software is software is really great. really great. really great. That's great, and That's great, and That's great, and actually, do you have actually, do you have actually, do you have any requests for any requests for any requests for the audience, people who the audience, people who the audience, people who use the M3 or the use the M3 or the use the M3 or the Minimax, or is there anything you'd Minimax, or is there anything you'd Minimax, or is there anything you'd like to get from like to get from like to get from them as feedback? Do them as feedback? Do them as feedback? Do you read, for example you read, for example , when people , when people , when people try to try to try to modify models modify models modify models or experiment or experiment , you know, make , you know, make , you know, make adjustments? Oh, what do adjustments? Oh, what do adjustments? Oh, what do you think is the best thing you think is the best thing you think is the best thing you can take from the you can take from the you can take from the community, for example, community, for example, community, for example, for future for future for future models? models? models? Mhm. I would say Mhm. I would say any problems that any problems that any problems that people face, people face, people face, especially with especially with especially with multimodality, multimodality, multimodality, right? This is the first right? This is the first right? This is the first time we've time we've time we've put all of this put all of this put all of this together. We will definitely be more together. We will definitely be more ambitious in this ambitious in this direction in the direction in the direction in the future. There future. There future. There may be some may be some may be some flaws right now, but we flaws right now, but we flaws right now, but we are working on are working on are working on fixing them. So, fixing them. So, fixing them. So, if there is feedback if there is feedback if there is feedback that the model is not working that the model is not working that the model is not working very well, we will very well, we will very well, we will definitely definitely definitely improve it in improve it in improve it in future versions. And future versions. And future versions. And also any also any also any features that features that features that people want. Let's say, people want. Let's say, people want. Let's say, for example, the effort to for example, the effort to for example, the effort to think, right? Some think, right? Some think, right? Some people ask for it.
-
people ask for it. people ask for it. Anyone can request it, Anyone can request it, Anyone can request it, and we will try to and we will try to and we will try to implement it in implement it in implement it in future models. future models. future models. So, are you seeing So, are you seeing So, are you seeing a lot of a lot of a lot of multimodality integration in multimodality integration in multimodality integration in the context of agents the context of agents the context of agents for writing code already? for writing code already? for writing code already? I think this is a I think this is a I think this is a bit of an unexplored bit of an unexplored bit of an unexplored area. area. That's true, that's true, but it That's true, that's true, but it can actually can actually can actually open up a lot of open up a lot of open up a lot of possibilities and a lot of possibilities and a lot of possibilities and a lot of agent applications agent applications . Let's say, for example, . Let's say, for example, . Let's say, for example, you want the model you want the model you want the model to read a PowerPoint or to read a PowerPoint or to read a PowerPoint or some reports that aren't some reports that aren't some reports that aren't very structured, and very structured, and very structured, and you want it to you want it to you want it to understand a very long understand a very long understand a very long video. Let's say you video. Let's say you video. Let's say you upload a long upload a long upload a long video and you want the video and you want the video and you want the model to act model to act model to act using certain using certain using certain tools after tools after tools after understanding it, this understanding it, this understanding it, this opens up a wide opens up a wide opens up a wide variety of agent use variety of agent use variety of agent use cases. cases. So, could the agent So, could the agent finally watch finally watch finally watch my YouTube video tutorial my YouTube video tutorial my YouTube video tutorial and understand how and understand how and understand how to use my to use my coding tools, as I coding tools, as I describe it, or something describe it, or something describe it, or something like that? like that? like that? Ha? Ha? Ha? Could an agent Could an agent Could an agent finally watch finally watch finally watch video tutorials on YouTube and video tutorials on YouTube and video tutorials on YouTube and understand things from them?
-
understand things from them? understand things from them? Yes, yes, yes. I think so Yes, yes, yes. I think so Yes, yes, yes. I think so . Do . Do . Do you use you use you use a lot of agent-based a lot of agent-based coding tools within your coding tools within your company? I company? I company? I mean, of course, for mean, of course, for mean, of course, for coding, but is coding, but is coding, but is it it it also used in research? also used in research? also used in research? Is this an automated Is this an automated Is this an automated part or not? How, how is part or not? How, how is part or not? How, how is that... that... that... Yes. Yes. Yes. working? working? working? We have our own We have our own We have our own research research research tools. We have tools. We have tools. We have created our own created our own created our own research research research tools that tools that tools that automate our automate our automate our workflows. I would workflows. I would workflows. I would say that many of say that many of say that many of our work our work our work processes are processes are processes are automated. You automated. You automated. You see how the see how the see how the most modern models most modern models most modern models are striving for are striving for are striving for features like features like features like kernel optimization, kernel optimization, kernel optimization, right? For example, a right? For example, a right? For example, a model that configures model that configures model that configures other models. other models. other models. Allow the model Allow the model Allow the model to create data. to create data. to create data. Autodata and all that. Autodata and all that. You can see that You can see that more and more models more and more models more and more models are capable of this, are capable of this, are capable of this, including the M3. including the M3. including the M3. In fact, we were In fact, we were In fact, we were very good at such very good at such very good at such cases. Longer cases. Longer cases. Longer horizons and horizons and horizons and kernel optimizations. So kernel optimizations. So kernel optimizations. So we can we can we can use use use the capabilities of models, the capabilities of models, the capabilities of models, combine them, and combine them, and combine them, and help in our help in our help in our daily work, daily work, daily work, making iterations even making iterations even making iterations even faster.
-
faster. Is M3 already building the M4? Is M3 already building the M4? We are building M 3.1. We are building M 3.1. We are building M 3.1. M 3.1, good. That's it, little M 3.1, good. That's it, little M 3.1, good. That's it, little pearl. pearl. pearl. Already. I'd Already. I'd like to close like to close with what you're excited about in the with what you're excited about in the with what you're excited about in the coming months, what coming months, what coming months, what you think could be you think could be you think could be game-changers in game-changers in game-changers in terms of features, or terms of features, or terms of features, or things you'd like to things you'd like to things you'd like to see in AI, or see in AI, or see in AI, or generally what you're generally what you're thinking about the most right now. I would say thinking about the most right now. I would say it's it's it's MGM. MGM. MGM. will happen. will happen. will happen. Many things are very Many things are very Many things are very exciting. But what exciting. But what exciting. But what I've been most I've been most fascinated with lately is fascinated with lately is multi-agents, which I multi-agents, which I multi-agents, which I think think think are used by many are used by many are used by many AI applications, AI applications, model routing, model routing, multi-agents—this multi-agents—this multi-agents—this opens up even more opens up even more opens up even more possibilities and possibilities and possibilities and allows you to perform more allows you to perform more allows you to perform more complex tasks. And complex tasks. And complex tasks. And it also tells us it also tells us it also tells us what the models are capable of and what they are what the models are capable of and what they are what the models are capable of and what they are not. And you can, you not. And you can, you not. And you can, you know, do a lot with know, do a lot with know, do a lot with it. This is it. This is it. This is quite fascinating. quite fascinating. quite fascinating. Thank you very much, Alife. We Thank you very much, Alife. We Thank you very much, Alife. We were glad to see you.
-
were glad to see you. were glad to see you. Thank you for inviting me. Thank you for inviting me. Thank you for inviting me. Thank you all.
Summary
The discussion focuses on leading open-source AI models, specifically highlighting MiniMax and its M3 release, a powerful model capable of understanding diverse data like video and images with a million-token context window. The key takeaway is that innovative AI companies like MiniMax are pushing the boundaries of model capabilities and accessibility.