← Back
AI Engineer July 31, 2026 19m

Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI

Read full transcript 9 segments
  1. Good morning everybody. Good morning everybody. My name is Ari Marcos. I'm the CEO and My name is Ari Marcos. I'm the CEO and My name is Ari Marcos. I'm the CEO and co-founder of Datlogy AI. Uh and really co-founder of Datlogy AI. Uh and really co-founder of Datlogy AI. Uh and really excited to kick off the data quality excited to kick off the data quality excited to kick off the data quality track uh today. Uh data quality is is track uh today. Uh data quality is is track uh today. Uh data quality is is what we live and track uh today. Uh data quality is is what we live and breathe uh at Datlogy. what we live and breathe uh at Datlogy. what we live and breathe uh at Datlogy. It's all we think about. In fact, the It's all we think about. In fact, the It's all we think about. In fact, the company's name literally means the company's name literally means the company's name literally means the science or study company's name literally means the science or study of data. Um, so very science or study of data. Um, so very science or study of data. Um, so very excited to see the increasing excitement excited to see the increasing excitement excited to see the increasing excitement and interest in this area excited to see the increasing excitement and interest in this area and a amazing and interest in this area and a amazing and interest in this area and a amazing lineup of talks uh today. So today I'm lineup of talks uh today. So today I'm lineup of talks uh today. So today I'm going to tell you about lineup of talks uh today. So today I'm going to tell you about why data quality going to tell you about why data quality going to tell you about why data quality is a compute multiplier that we're all is a compute multiplier that we're all is a compute multiplier that we're all overlooking and where we can make overlooking and where we can make overlooking and where we can make massive gains overlooking and where we can make massive gains by just working on better massive gains by just working on better massive gains by just working on better data. Um, we've seen that compute data. Um, we've seen that compute data. Um, we've seen that compute availability data. Um, we've seen that compute availability over the last six months availability over the last six months availability over the last six months has become extremely scarce and is only has become extremely scarce and is only has become extremely scarce and is only getting worse. has become extremely scarce and is only getting worse. We saw H100 prices getting worse. We saw H100 prices getting worse. We saw H100 prices reverse their several year-long drop reverse their several year-long drop reverse their several year-long drop which is normal for hardware reverse their several year-long drop which is normal for hardware and all of which is normal for hardware and all of which is normal for hardware and all of a sudden come up where now they're about a sudden come up where now they're about a sudden come up where now they're about 40% up from their lows um at the end of 40% up from their lows um at the end of 40% up from their lows um at the end of last year. 40% up from their lows um at the end of last year. Um as test time compute has last year. Um as test time compute has last year. Um as test time compute has become a critical part of models and as become a critical part of models and as become a critical part of models and as we become a critical part of models and as we have put more and more thinking we have put more and more thinking we have put more and more thinking tokens in we're seeing that the number tokens in we're seeing that the number tokens in we're seeing that the number the token usage is absolutely the token usage is absolutely the token usage is absolutely skyrocketing. Um reasoning models use skyrocketing. Um reasoning models use skyrocketing. Um reasoning models use eight times as many tokens um as eight times as many tokens um as eight times as many tokens um as non-reasoning eight times as many tokens um as non-reasoning models and that's non-reasoning models and that's non-reasoning models and that's projected to 5x again in the next year projected to 5x again in the next year projected to 5x again in the next year or so. Um, so the number of tokens we're or so. Um, so the number of tokens we're or so. Um, so the number of tokens we're pushing through goes higher and higher pushing through goes higher and higher pushing through goes higher and higher and that constrains compute even and that constrains compute even and that constrains compute even further. And this has led actually to a further. And this has led actually to a further. And this has led actually to a world where it's not implausible today world where it's not implausible today world where it's not implausible today that we might see access to some of the that we might see access to some of the that we might see access to some of the frontier APIs actually get limited that we might see access to some of the frontier APIs actually get limited or go frontier APIs actually get limited or go frontier APIs actually get limited or go away. Um, as one example, Google just away. Um, as one example, Google just away. Um, as one example, Google just capped uh Meta's Gemini usage because of capped uh Meta's Gemini usage because of capped uh Meta's Gemini usage because of inference constraints. OpenAI has inference constraints. OpenAI has inference constraints. OpenAI has effectively started selling token effectively started selling token effectively started selling token futures or you can guarantee token futures or you can guarantee token futures or you can guarantee token capacity some amount of time into the

  2. capacity some amount of time into the capacity some amount of time into the future. This is only necessary because future. This is only necessary because future. This is only necessary because they are legitimately future. This is only necessary because they are legitimately wondering there they are legitimately wondering there they are legitimately wondering there might be a world where access to might be a world where access to might be a world where access to frontier API tokens is limited not as a frontier API tokens is limited not as a frontier API tokens is limited not as a business decision but because there's business decision but because there's business decision but because there's just simply not enough inference and just simply not enough inference and just simply not enough inference and first party products will be first party products will be first party products will be prioritized. So first party products will be prioritized. So in a world where compute prioritized. So in a world where compute prioritized. So in a world where compute is increasingly scarce and you need to is increasingly scarce and you need to is increasingly scarce and you need to make models better. Well, what do you make models better. Well, what do you make models better. Well, what do you do? make models better. Well, what do you do? Well, we work on data. Um I you're do? Well, we work on data. Um I you're do? Well, we work on data. Um I you're going to hear me say this over and over going to hear me say this over and over going to hear me say this over and over again. Data quality is a compute again. Data quality is a compute again. Data quality is a compute multiplier again. Data quality is a compute multiplier because what it does is it it multiplier because what it does is it it multiplier because what it does is it it makes the learning curve steeper. So makes the learning curve steeper. So makes the learning curve steeper. So this is just a very simple schematic makes the learning curve steeper. So this is just a very simple schematic of this is just a very simple schematic of this is just a very simple schematic of performance on the y-axis as a function performance on the y-axis as a function performance on the y-axis as a function of data on the x-axis. Note that this of data on the x-axis. Note that this of data on the x-axis. Note that this x-axis, you of data on the x-axis. Note that this x-axis, you could swap it out for data, x-axis, you could swap it out for data, x-axis, you could swap it out for data, for compute, for time, for dollars. for compute, for time, for dollars. for compute, for time, for dollars. They're all the same x-axis They're all the same x-axis They're all the same x-axis fundamentally. Um, and if you can make fundamentally. Um, and if you can make fundamentally. Um, and if you can make data quality better, you can turn this data quality better, you can turn this data quality better, you can turn this gray data quality better, you can turn this gray curve um into this blue curve. And gray curve um into this blue curve. And gray curve um into this blue curve. And that means that now you can get that means that now you can get that means that now you can get dramatically better performance that means that now you can get dramatically better performance for the dramatically better performance for the dramatically better performance for the same compute budget as if you had same compute budget as if you had same compute budget as if you had trained with far more compute. Um and trained with far more compute. Um and trained with far more compute. Um and similarly you can get the same similarly you can get the same similarly you can get the same performance for a much smaller compute performance for a much smaller compute performance for a much smaller compute budget. Um which performance for a much smaller compute budget. Um which is exactly showing you budget. Um which is exactly showing you budget. Um which is exactly showing you how you can get performance as if you how you can get performance as if you how you can get performance as if you had spent 10 times this how you can get performance as if you had spent 10 times this 10 times as much had spent 10 times this 10 times as much had spent 10 times this 10 times as much on compute and really shows this compute on compute and really shows this compute on compute and really shows this compute multiplier point.

  3. on compute and really shows this compute multiplier point. Um so how do you multiplier point. Um so how do you multiplier point. Um so how do you actually do this? Well fundamentally the actually do this? Well fundamentally the actually do this? Well fundamentally the idea is we want to make it so that we idea is we want to make it so that we idea is we want to make it so that we get the maximum signal per idea is we want to make it so that we get the maximum signal per token and per get the maximum signal per token and per get the maximum signal per token and per batch. Um for those that are a little batch. Um for those that are a little batch. Um for those that are a little bit more technically what we want to bit more technically what we want to bit more technically what we want to want to do here is bit more technically what we want to want to do here is maximize the marginal want to do here is maximize the marginal want to do here is maximize the marginal information gain per data point when we information gain per data point when we information gain per data point when we show it to the model. what data is going show it to the model. what data is going show it to the model. what data is going to teach show it to the model. what data is going to teach the model the most. Um, and to teach the model the most. Um, and to teach the model the most. Um, and that's all about finding data that's that's all about finding data that's that's all about finding data that's relevant to that's all about finding data that's relevant to the use cases that you want. relevant to the use cases that you want. relevant to the use cases that you want. One thing that you know is very One thing that you know is very One thing that you know is very important is that there's no one golden important is that there's no one golden important is that there's no one golden data set important is that there's no one golden data set to rule them all that's good data set to rule them all that's good data set to rule them all that's good for everything no matter what you want for everything no matter what you want for everything no matter what you want to do. A data set's only going to be to do. A data set's only going to be to do. A data set's only going to be optimal with respect to do. A data set's only going to be optimal with respect to a particular set optimal with respect to a particular set optimal with respect to a particular set of output tasks that you want the model of output tasks that you want the model of output tasks that you want the model to do. So if you want a great legal to do. So if you want a great legal to do. So if you want a great legal model, to do. So if you want a great legal model, you're going to want legal data model, you're going to want legal data model, you're going to want legal data more than healthcare data and vice more than healthcare data and vice more than healthcare data and vice versa. It needs to be diverse. A more than healthcare data and vice versa. It needs to be diverse. A lot of versa. It needs to be diverse. A lot of versa. It needs to be diverse. A lot of the issues we see with model robustness the issues we see with model robustness the issues we see with model robustness and brittleleness comes from try and brittleleness comes from try and brittleleness comes from try training on data that's not diverse training on data that's not diverse training on data that's not diverse enough. training on data that's not diverse enough. So the model can answer a enough. So the model can answer a enough. So the model can answer a question correctly if it's presented question correctly if it's presented question correctly if it's presented just so but if it's presented a little just so but if it's presented a little just so but if it's presented a little bit differently now everything breaks um bit differently now everything breaks um bit differently now everything breaks um needs to be information and you have to needs to be information and you have to needs to be information and you have to mix the data correctly. This is a needs to be information and you have to mix the data correctly. This is a hugely mix the data correctly. This is a hugely mix the data correctly. This is a hugely difficult part of you now have many difficult part of you now have many difficult part of you now have many different sources. How do you combine different sources. How do you combine different sources. How do you combine them to actually drive the largest them to actually drive the largest them to actually drive the largest improvement in them to actually drive the largest improvement in performance?

  4. improvement in performance? improvement in performance? So um this is a high level of what we do So um this is a high level of what we do So um this is a high level of what we do at here. So um this is a high level of what we do at here. Um you can think of us as the at here. Um you can think of us as the at here. Um you can think of us as the oil refinery for data. We don't source oil refinery for data. We don't source oil refinery for data. We don't source new tokens like many data providers. new tokens like many data providers. new tokens like many data providers. Rather we take existing new tokens like many data providers. Rather we take existing tokens coming Rather we take existing tokens coming Rather we take existing tokens coming from public data sets, proprietary data from public data sets, proprietary data from public data sets, proprietary data sets and licensed data sets and make sets and licensed data sets and make sets and licensed data sets and make them way sets and licensed data sets and make them way better. And how do we do that? them way better. And how do we do that? them way better. And how do we do that? Well, we do that through these four C's. Well, we do that through these four C's. Well, we do that through these four C's. Um, clean, curate, create, and Well, we do that through these four C's. Um, clean, curate, create, and compose. Um, clean, curate, create, and compose. Um, clean, curate, create, and compose. Um, so cleaning is fairly Um, so cleaning is fairly Um, so cleaning is fairly straightforward. This is doing things straightforward. This is doing things straightforward. This is doing things like heruristic filters, all a gopher like heruristic filters, all a gopher like heruristic filters, all a gopher and things like that. Removing documents and things like that. Removing documents and things like that. Removing documents that have, you know, only 10 characters that have, you know, only 10 characters that have, you know, only 10 characters in them or all winging. that have, you know, only 10 characters in them or all winging. That's kind of in them or all winging. That's kind of in them or all winging. That's kind of basic table stakes. Um, benchmark basic table stakes. Um, benchmark basic table stakes. Um, benchmark decontamination is incredibly important. decontamination is incredibly important. decontamination is incredibly important. As I'm sure you all know, benchmaxing As I'm sure you all know, benchmaxing As I'm sure you all know, benchmaxing has become a real problem and makes it has become a real problem and makes it has become a real problem and makes it very difficult to interpret model very difficult to interpret model very difficult to interpret model results. So, we rigorously decontaminate results. So, we rigorously decontaminate results. So, we rigorously decontaminate all of our training data with respect to all of our training data with respect to all of our training data with respect to all downstream all of our training data with respect to all downstream benchmarks with a pretty all downstream benchmarks with a pretty all downstream benchmarks with a pretty low engram to ensure that that's the low engram to ensure that that's the low engram to ensure that that's the case. That gets you to a point where now case. That gets you to a point where now case. That gets you to a point where now you can feed case. That gets you to a point where now you can feed the model into the data, you can feed the model into the data, you can feed the model into the data, but it's still sorry, feed the data into but it's still sorry, feed the data into but it's still sorry, feed the data into the model, but it's still not very good. the model, but it's still not very good. the model, but it's still not very good. So then how do you make it better? the model, but it's still not very good. So then how do you make it better? Well, So then how do you make it better? Well, So then how do you make it better? Well, it's a combination of many things it's a combination of many things it's a combination of many things ranging from quality classifiers and ranging from quality classifiers and ranging from quality classifiers and tonomy across different topics ranging from quality classifiers and tonomy across different topics and tonomy across different topics and tonomy across different topics and balancing that redundancy reduction. So balancing that redundancy reduction. So balancing that redundancy reduction. So removing data points that are not the removing data points that are not the removing data points that are not the same that are semantically removing data points that are not the same that are semantically similar but same that are semantically similar but same that are semantically similar but convey very similar information even if convey very similar information even if convey very similar information even if they're not the same pixels themselves they're not the same pixels themselves they're not the same pixels themselves say they're not the same pixels themselves say upsampling and downsampling data say upsampling and downsampling data say upsampling and downsampling data points based off the quality and the points based off the quality and the points based off the quality and the relevance and then task distribution relevance and then task distribution relevance and then task distribution matching identifying what data do you matching identifying what data do you matching identifying what data do you actually need in order to solve this actually need in order to solve this actually need in order to solve this given task. That now gives you actually need in order to solve this given task. That now gives you a data given task. That now gives you a data given task. That now gives you a data set that is very high quality but is set that is very high quality but is set that is very high quality but is typically still too small. And that's typically still too small. And that's typically still too small. And that's where synthetic typically still too small. And that's where synthetic data comes in. Now we where synthetic data comes in. Now we where synthetic data comes in. Now we can go and rephrase that data as can go and rephrase that data as can go and rephrase that data as effectively a very fancy form of data effectively a very fancy form of data effectively a very fancy form of data augmentation effectively a very fancy form of data augmentation to produce dramatically augmentation to produce dramatically augmentation to produce dramatically more data in many different formats and

  5. more data in many different formats and more data in many different formats and this helps a lot more data in many different formats and this helps a lot both with kind of uh this helps a lot both with kind of uh this helps a lot both with kind of uh data size and with diversity um because data size and with diversity um because data size and with diversity um because we can really inject a lot data size and with diversity um because we can really inject a lot of diversity we can really inject a lot of diversity we can really inject a lot of diversity in through this and then finally you in through this and then finally you in through this and then finally you know have these data sets how do you know have these data sets how do you know have these data sets how do you combine them and how do you com and know have these data sets how do you combine them and how do you com and how combine them and how do you com and how combine them and how do you com and how do you sequence them across different do you sequence them across different do you sequence them across different training stages it's now become table training stages it's now become table training stages it's now become table stakes that any large model is generally stakes that any large model is generally stakes that any large model is generally trained for at least three phases of trained for at least three phases of trained for at least three phases of data um how do you do that um and can data um how do you do that um and can data um how do you do that um and can you actually even do continuous data um how do you do that um and can you actually even do continuous uh you actually even do continuous uh you actually even do continuous uh curricula and things like that which is curricula and things like that which is curricula and things like that which is a lot of what we work on at Dtology. Um a lot of what we work on at Dtology. Um a lot of what we work on at Dtology. Um and that ultimately a lot of what we work on at Dtology. Um and that ultimately gets you a much and that ultimately gets you a much and that ultimately gets you a much better data set out. All right, so better data set out. All right, so better data set out. All right, so that's a high level of kind of what we that's a high level of kind of what we that's a high level of kind of what we need to do. What can you that's a high level of kind of what we need to do. What can you actually get need to do. What can you actually get need to do. What can you actually get out of this? Can this actually really out of this? Can this actually really out of this? Can this actually really make a massive difference? Um so we've make a massive difference? Um so we've make a massive difference? Um so we've about half of our team at Datlogy are about half of our team at Datlogy are about half of our team at Datlogy are just researchers. Um and you know we uh just researchers. Um and you know we uh just researchers. Um and you know we uh do all of our own research just researchers. Um and you know we uh do all of our own research on how we do do all of our own research on how we do do all of our own research on how we do data curation effectively. um because data curation effectively. um because data curation effectively. um because this is such a critical part of the uh this is such a critical part of the uh this is such a critical part of the uh model building pipeline uh there's very model building pipeline uh there's very model building pipeline uh there's very little published here um because there's little published here um because there's little published here um because there's a very strong disincentive not little published here um because there's a very strong disincentive not to share a very strong disincentive not to share a very strong disincentive not to share how you do this um the kind of how you do this um the kind of how you do this um the kind of foundational paper for daty was how you do this um the kind of foundational paper for daty was one I foundational paper for daty was one I foundational paper for daty was one I wrote when I was at meta called beyond wrote when I was at meta called beyond wrote when I was at meta called beyond scaling laws um which was fortunate to scaling laws um which was fortunate to scaling laws um which was fortunate to get a best paper at nurips a scaling laws um which was fortunate to get a best paper at nurips a couple get a best paper at nurips a couple get a best paper at nurips a couple years ago which showed that if you years ago which showed that if you years ago which showed that if you choose your data correctly you can choose your data correctly you can choose your data correctly you can actually bend the scaling laws itself actually bend the scaling laws itself actually bend the scaling laws itself you can change the exponent um and you can change the exponent um and you can change the exponent um and that's because you're now not wasting that's because you're now not wasting that's because you're now not wasting your time looking at redundant that's because you're now not wasting your time looking at redundant or your time looking at redundant or your time looking at redundant or unnecessary data That was very much the unnecessary data That was very much the unnecessary data That was very much the proof of principle for all of Dtology.

  6. proof of principle for all of Dtology. proof of principle for all of Dtology. And we've since proof of principle for all of Dtology. And we've since expanded this into many And we've since expanded this into many And we've since expanded this into many public research releases. We've shared public research releases. We've shared public research releases. We've shared of various ways to improve models just of various ways to improve models just of various ways to improve models just through of various ways to improve models just through data curation. I'm going to go through data curation. I'm going to go through data curation. I'm going to go through a couple of those results now um through a couple of those results now um through a couple of those results now um and show you what we've been able to to and show you what we've been able to to and show you what we've been able to to to achieve. and show you what we've been able to to to achieve. Um so first let's talk about to achieve. Um so first let's talk about to achieve. Um so first let's talk about vision language models. How can we vision language models. How can we vision language models. How can we improve vision language models. How can we improve um VLMs just through data improve um VLMs just through data improve um VLMs just through data curation alone? Um so in this case what curation alone? Um so in this case what curation alone? Um so in this case what we curation alone? Um so in this case what we did is we took the mammoth data set. we did is we took the mammoth data set. we did is we took the mammoth data set. This is a fairly small data set about 25 This is a fairly small data set about 25 This is a fairly small data set about 25 billion um This is a fairly small data set about 25 billion um tokens uh that we use for the billion um tokens uh that we use for the billion um tokens uh that we use for the purposes of training the fusion adapter purposes of training the fusion adapter purposes of training the fusion adapter layer purposes of training the fusion adapter layer between your um text uh your text layer between your um text uh your text layer between your um text uh your text your text model and your vision model. your text model and your vision model. your text model and your vision model. Um and what you can see this is a Um and what you can see this is a Um and what you can see this is a scaling plot where we have error on the scaling plot where we have error on the scaling plot where we have error on the y-axis as a function of log scaling plot where we have error on the y-axis as a function of log flops um on y-axis as a function of log flops um on y-axis as a function of log flops um on the x-axis. Um so there's about a scale the x-axis. Um so there's about a scale the x-axis. Um so there's about a scale of a thousand from the left the x-axis. Um so there's about a scale of a thousand from the left uh most part of a thousand from the left uh most part of a thousand from the left uh most part of this plot to the rightmost part. Um of this plot to the rightmost part. Um of this plot to the rightmost part. Um and what you can see is that if you look and what you can see is that if you look and what you can see is that if you look at the parto frontier defined and what you can see is that if you look at the parto frontier defined by many of at the parto frontier defined by many of at the parto frontier defined by many of the best uh public VLMs like the Quen 3 the best uh public VLMs like the Quen 3 the best uh public VLMs like the Quen 3 series and 3.5 intern VL the best uh public VLMs like the Quen 3 series and 3.5 intern VL etc you can see series and 3.5 intern VL etc you can see series and 3.5 intern VL etc you can see that a model trained on daties data are that a model trained on daties data are that a model trained on daties data are able to go well well beyond that a model trained on daties data are able to go well well beyond uh that able to go well well beyond uh that able to go well well beyond uh that frontier and I'll note this is actually frontier and I'll note this is actually frontier and I'll note this is actually without any post- training as well. Uh without any post- training as well. Uh without any post- training as well. Uh so you can get without any post- training as well. Uh so you can get very strong performance so you can get very strong performance so you can get very strong performance across many different benchmarks. Um so across many different benchmarks. Um so across many different benchmarks. Um so just looking at kind of taking the input just looking at kind of taking the input just looking at kind of taking the input data set that gray diamond there that's data set that gray diamond there that's data set that gray diamond there that's the input data set we use to do our the input data set we use to do our the input data set we use to do our curation. Um you can see the input data set we use to do our curation. Um you can see that just curation. Um you can see that just curation. Um you can see that just through curation you're able to get through curation you're able to get through curation you're able to get around a 14 absolute percentage point around a 14 absolute percentage point around a 14 absolute percentage point improvement um holding around a 14 absolute percentage point improvement um holding everything else improvement um holding everything else improvement um holding everything else constant just through better data alone.

  7. constant just through better data alone. constant just through better data alone. Um and not only that you can also see Um and not only that you can also see Um and not only that you can also see that we Um and not only that you can also see that we can roughly match the that we can roughly match the that we can roughly match the performance of quen 3.54b come with performance of quen 3.54b come with performance of quen 3.54b come with about a percentage of it um while using about a percentage of it um while using about a percentage of it um while using 145x less training compute in a world 145x less training compute in a world 145x less training compute in a world with no with less compute. 145x less training compute in a world with no with less compute. How do you do with no with less compute. How do you do with no with less compute. How do you do more? You make data better and now it's more? You make data better and now it's more? You make data better and now it's as if you had 100 times um the compute. as if you had 100 times um the compute. as if you had 100 times um the compute. Um, interestingly I mentioned that Um, interestingly I mentioned that Um, interestingly I mentioned that reasoning models are be are using tokens reasoning models are be are using tokens reasoning models are be are using tokens at a very high rate as well. reasoning models are be are using tokens at a very high rate as well. Well, at a very high rate as well. Well, at a very high rate as well. Well, another thing that we found is that data another thing that we found is that data another thing that we found is that data curation can also lead uh to more curation can also lead uh to more curation can also lead uh to more concise answers curation can also lead uh to more concise answers um depending on how you concise answers um depending on how you concise answers um depending on how you represent the data. So what's plotted represent the data. So what's plotted represent the data. So what's plotted here is the mean number of tokens represent the data. So what's plotted here is the mean number of tokens per here is the mean number of tokens per here is the mean number of tokens per response um across all the same set of response um across all the same set of response um across all the same set of models for the large most part um that models for the large most part um that models for the large most part um that we just showed. models for the large most part um that we just showed. Um and you can see that we just showed. Um and you can see that we just showed. Um and you can see that that models trained on data those three that models trained on data those three that models trained on data those three blue lines right at the top are all blue lines right at the top are all blue lines right at the top are all extremely concise. Um and if we do kind extremely concise. Um and if we do kind extremely concise. Um and if we do kind of the same sort of plot um but now on of the same sort of plot um but now on of the same sort of plot um but now on the x-axis instead of of the same sort of plot um but now on the x-axis instead of log training flops the x-axis instead of log training flops the x-axis instead of log training flops this is now um log flops per response. this is now um log flops per response. this is now um log flops per response. Um so this is inference efficiency. this is now um log flops per response. Um so this is inference efficiency. Um Um so this is inference efficiency. Um Um so this is inference efficiency. Um you can still see that we go well on the you can still see that we go well on the you can still see that we go well on the ex by beating that router frontier um ex by beating that router frontier um ex by beating that router frontier um and you know roughly get uh similar and you know roughly get uh similar and you know roughly get uh similar performance to quen 35 with 35 fewer uh performance to quen 35 with 35 fewer uh performance to quen 35 with 35 fewer uh times fewer performance to quen 35 with 35 fewer uh times fewer flops uh per correct answer.

  8. times fewer flops uh per correct answer. times fewer flops uh per correct answer. So data curation can make a huge impact So data curation can make a huge impact So data curation can make a huge impact in VLMs. What about So data curation can make a huge impact in VLMs. What about text models? Um, one in VLMs. What about text models? Um, one in VLMs. What about text models? Um, one of the most uh challenging things about of the most uh challenging things about of the most uh challenging things about many models is that they work very well many models is that they work very well many models is that they work very well on English many models is that they work very well on English data, but they don't work on English data, but they don't work on English data, but they don't work well um for non-western use cases in well um for non-western use cases in well um for non-western use cases in general. The internet well um for non-western use cases in general. The internet is an extremely general. The internet is an extremely general. The internet is an extremely biased view of the world that does not biased view of the world that does not biased view of the world that does not represent the world uniformly at all. represent the world uniformly at all. represent the world uniformly at all. And this has represent the world uniformly at all. And this has major implications for And this has major implications for And this has major implications for fairness and for the usability of these fairness and for the usability of these fairness and for the usability of these models across the world. I don't want fairness and for the usability of these models across the world. I don't want to models across the world. I don't want to models across the world. I don't want to live in a future where only developed live in a future where only developed live in a future where only developed countries can access this very countries can access this very countries can access this very effectively. Um, so how do you do countries can access this very effectively. Um, so how do you do this? effectively. Um, so how do you do this? effectively. Um, so how do you do this? Well, curation again can be a massive Well, curation again can be a massive Well, curation again can be a massive lever here. Um so what I'm plotting here lever here. Um so what I'm plotting here lever here. Um so what I'm plotting here now is lever here. Um so what I'm plotting here now is a similar plot error on the y- now is a similar plot error on the y- now is a similar plot error on the y- axis as a function of log flops about axis as a function of log flops about axis as a function of log flops about 100x going from left to right here. axis as a function of log flops about 100x going from left to right here. Um 100x going from left to right here. Um 100x going from left to right here. Um this is highlighting multilingual MMLU this is highlighting multilingual MMLU this is highlighting multilingual MMLU performance. You can see we have a parto performance. You can see we have a parto performance. You can see we have a parto frontier here performance. You can see we have a parto frontier here defined by many models. frontier here defined by many models. frontier here defined by many models. The quen models um the liquid some of The quen models um the liquid some of The quen models um the liquid some of the liquid models. Um the The quen models um the liquid some of the liquid models. Um the green square the liquid models. Um the green square the liquid models. Um the green square is tiny coher's best multilingual model. is tiny coher's best multilingual model. is tiny coher's best multilingual model. You can see again that we're well off You can see again that we're well off You can see again that we're well off the predto You can see again that we're well off the predto frontier with a couple things the predto frontier with a couple things the predto frontier with a couple things I really want to highlight. First off we I really want to highlight. First off we I really want to highlight. First off we only use 8% of the data here as only use 8% of the data here as only use 8% of the data here as multilingual only use 8% of the data here as multilingual tokens. Um, so most multilingual tokens. Um, so most multilingual tokens. Um, so most languages actually only had um at max 6 languages actually only had um at max 6 languages actually only had um at max 6 billion languages actually only had um at max 6 billion tokens here. So these are not billion tokens here. So these are not billion tokens here. So these are not massive amounts of data uh in the massive amounts of data uh in the massive amounts of data uh in the non-English languages that are going in non-English languages that are going in non-English languages that are going in here. Um you can again see we get the here. Um you can again see we get the here. Um you can again see we get the same sort of compute multiplier effect.

  9. same sort of compute multiplier effect. same sort of compute multiplier effect. We're a little better than Quen same sort of compute multiplier effect. We're a little better than Quen 3 um We're a little better than Quen 3 um We're a little better than Quen 3 um while while having roughly 8x less while while having roughly 8x less while while having roughly 8x less compute budget um here. So you can make compute budget um here. So you can make compute budget um here. So you can make a

Summary

The main theme is that data quality is an overlooked compute multiplier, with significant implications for AI model performance and resource availability. The discussion highlights the increasing scarcity and cost of compute, citing rising H100 prices and the burgeoning token usage in reasoning models. The practical takeaway is that focusing on better data quality can unlock massive gains in compute efficiency, preventing potential limitations in API access and prioritizing first-party products.

View original episode ↗