← Back
AI Engineer October 10, 2026 17m

From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model

Read full transcript 13 segments
  1. Hello everyone. Thank you for Hello everyone. Thank you for joining me. joining me. Um, yes, today I'm going to Um, yes, today I'm going to talk about Ray- talk about Ray- actors, visual actors, visual actors, visual tokens, and GIL. Um, tokens, and GIL. Um, tokens, and GIL. Um, developing a developing a developing a multimodal multimodal multimodal SFT data pipeline SFT data pipeline SFT data pipeline that keeps GPUs that keeps GPUs that keeps GPUs busy. So, there's a busy. So, there's a lot of talk about lot of talk about , you know, how to , you know, how to , you know, how to optimize optimize optimize learning paradigms. learning paradigms. But I want to provide a But I want to provide a different perspective on the entire different perspective on the entire different perspective on the entire ecosystem of what it ecosystem of what it ecosystem of what it takes to takes to takes to train a model. And train a model. And train a model. And so the question I'm so the question I'm so the question I'm asking here is: is asking here is: is asking here is: is your learning cycle really your learning cycle really your learning cycle really GPU-dependent right now? GPU-dependent right now? GPU-dependent right now? Is it GPU-locked? Is it GPU-locked? Is it GPU-locked? I think we'll learn I think we'll learn I think we'll learn and see in this and see in this and see in this experiment that experiment that experiment that the hardest part of the hardest part of the hardest part of optimizing optimizing multimodal multimodal learning throughput isn't learning throughput isn't learning throughput isn't necessarily necessarily necessarily optimizing the kernels. This is optimizing the kernels. This is optimizing the kernels. This is to ensure that to ensure that to ensure that the GPUs are constantly the GPUs are constantly the GPUs are constantly receiving data. So receiving data. So , let's take a look at an , let's take a look at an , let's take a look at an overview of this topic, overview of this topic, overview of this topic, what this what this what this report will be about. We will create a report will be about. We will create a report will be about. We will create a basic basic basic multimodal multimodal multimodal learning pipeline. Um, learning pipeline. Um, learning pipeline. Um, and then we'll run into and then we'll run into and then we'll run into bottlenecks and bottlenecks and bottlenecks and tackle them one by tackle them one by tackle them one by one. So, one. So, one. So, conditionally it is: one is conditionally it is: one is parallelism. How quickly parallelism. How quickly parallelism. How quickly can you process a can you process a can you process a single image single image single image before feeding before feeding before feeding it to the model? Two— it to the model? Two— preload. How do preload. How do we make sure we have the we make sure we have the we make sure we have the data we are requesting?

  2. data we are requesting? Three—how to minimize Three—how to minimize data copying between data copying between data copying between different processes? And different processes? And different processes? And four—other four—other four—other additional additional framework optimizations. Um, framework optimizations. Um, just setting up just setting up just setting up all the little things. Um, yeah all the little things. Um, yeah , let's define , let's define , let's define our source data. We our source data. We our source data. We all know that GPUs are all know that GPUs are all know that GPUs are expensive, and we expensive, and we expensive, and we usually want to usually want to usually want to maximize their maximize their maximize their use while we have them use while we have them use while we have them . . Multimodal Multimodal processing processing processing differs from differs from differs from text processing because there are text processing because there are text processing because there are many many many CPU-intensive CPU-intensive operations that must operations that must operations that must occur before occur before occur before the data is fed into the the data is fed into the the data is fed into the model. This could be, model. This could be, model. This could be, you know, loading, you know, loading, you know, loading, resizing, or resizing, or resizing, or normalizing normalizing normalizing images images images using libraries using libraries using libraries like PIL. For audio, you like PIL. For audio, you like PIL. For audio, you do a lot of do a lot of do a lot of downloading and downloading and downloading and trimming. You can trimming. You can trimming. You can use use use NumPy arrays. And video NumPy arrays. And video is a combination of is a combination of is a combination of both. Um, and it's also both. Um, and it's also both. Um, and it's also true that there are true that there are true that there are usually usually many more CPUs many more CPUs than GPUs in your training clusters. However, the GPU is what is than GPUs in your training clusters. However, the GPU is what is than GPUs in your training clusters. However, the GPU is what is starving in this starving in this starving in this pipeline. So pipeline. So pipeline. So today we're going to today we're going to today we're going to introduce an equation introduce an equation called the called the latency factor, which I latency factor, which I generally phrase as the generally phrase as the generally phrase as the time it takes to prepare the time it takes to prepare the time it takes to prepare the data and transfer it data and transfer it data and transfer it to the GPU. Um, and then we to the GPU. Um, and then we to the GPU. Um, and then we divide all of that by the time it divide all of that by the time it divide all of that by the time it takes for both of takes for both of takes for both of those operations, plus the those operations, plus the those operations, plus the training itself. So, training itself. So, training itself. So, essentially, the denominator essentially, the denominator is your full pipeline is your full pipeline . So, we'll see that

  3. . So, we'll see that . So, we'll see that during our during our during our launch, we'll go launch, we'll go launch, we'll go from 15% load to from 15% load to from 15% load to 90% load. So, 90% load. So, 90% load. So, our baseline our baseline our baseline spends about 85% of spends about 85% of its time its time its time retrieving data, while retrieving data, while retrieving data, while after optimization it spends less than after optimization it spends less than 10%. So within the scope of 10%. So within the scope of 10%. So within the scope of this report, this report, this report, graphics processing units ( graphics processing units ( GPUs) are not the main GPUs) are not the main GPUs) are not the main focus. The goal is to focus. The goal is to focus. The goal is to ensure ensure that data never that data never that data never becomes a “ becomes a “ becomes a “ bottleneck.” Here's a little bottleneck.” Here's a little bottleneck.” Here's a little information about what information about what information about what our stack will look like. our stack will look like. We take the base We take the base Qwen-3-VL model and Qwen-3-VL model and Qwen-3-VL model and try to train try to train try to train it on the ChatVQA it on the ChatVQA it on the ChatVQA and multi-pass and multi-pass and multi-pass VLM dialogue paradigms. So, VLM dialogue paradigms. So, VLM dialogue paradigms. So, to summarize: before to summarize: before the GPUs start the GPUs start the GPUs start computing, we computing, we computing, we need to load need to load need to load the data, transform it, the data, transform it, the data, transform it, and then pass it to the GPU and then pass it to the GPU . Okay, that's an overview of what . Okay, that's an overview of what our data actually looks like.

  4. We will We will use use use JSONL files, and JSONL files, and JSONL files, and the images are linked the images are linked the images are linked relative to the location of relative to the location of the example. Let's assume the example. Let's assume our data is simply our data is simply our data is simply stored in S3. stored in S3. For example, For example, the image 33471.jpeg appears to be the image 33471.jpeg appears to be a photo of a a photo of a a photo of a blue bus. Yes blue bus. Yes , this is a layout of our , this is a layout of our , this is a layout of our system. On the left we have the system. On the left we have the system. On the left we have the data generator, which is data generator, which is data generator, which is responsible for responsible for responsible for loading, loading, loading, converting, and converting, and converting, and transmitting them. This is transmitting them. This is transmitting them. This is denoted as denoted as denoted as fan out to fan out to Megatron data parallelism ranks. Parallel ranks Parallel ranks receive parts of receive parts of receive parts of the data determined the data determined the data determined by the coordinator. The data by the coordinator. The data by the coordinator. The data is compared, after is compared, after is compared, after which a which a which a forward and reverse forward and reverse forward and reverse pass is performed. Megatron pass is performed. Megatron pass is performed. Megatron takes care of takes care of takes care of optimizing the optimizing the optimizing the model weights between DP ranks, model weights between DP ranks, model weights between DP ranks, and we will assume that and we will assume that and we will assume that it works as a “black it works as a “black it works as a “black box”. So, box”. So, box”. So, let's start with a naive let's start with a naive let's start with a naive basic path. For basic path. For basic path. For each example, we each example, we each example, we retrieve an image retrieve an image retrieve an image from the object store.

  5. from the object store. We tokenize them, We tokenize them, resize them, resize them, resize them, normalize them, and then normalize them, and then normalize them, and then feed them into the model. feed them into the model. When we do this, When we do this, we see the sequential we see the sequential we see the sequential operations marked in operations marked in operations marked in red: retrieving red: retrieving red: retrieving all the images one by all the images one by all the images one by one and processing them in turn one and processing them in turn one and processing them in turn . So if your . So if your . So if your training process training process training process requests a requests a requests a batch size of 256, you will batch size of 256, you will batch size of 256, you will have to have to have to process process process all those images sequentially. And all those images sequentially. And all those images sequentially. And then we see that our then we see that our then we see that our learning time learning time learning time is between 15 and 20 is between 15 and 20 is between 15 and 20 seconds. However, the time seconds. However, the time seconds. However, the time spent spent spent directly on directly on directly on data acquisition data acquisition data acquisition exceeds 100 seconds. exceeds 100 seconds. exceeds 100 seconds. Therefore, we Therefore, we Therefore, we only get 15% utilization of only get 15% utilization of only get 15% utilization of our GPUs. You can also see in the our GPUs. You can also see in the our GPUs. You can also see in the profiling graph profiling graph profiling graph that we that we that we spend most of our spend most of our spend most of our time downloading time downloading time downloading from S3. And at the end, from S3. And at the end, from S3. And at the end, you can see a small " you can see a small " tail" associated with tail" associated with tail" associated with loading the loading the loading the images themselves through the images themselves through the images themselves through the PIL library. So, PIL library. So, PIL library. So, okay. Solution one: okay. Solution one: okay. Solution one: let's try to add let's try to add let's try to add parallelism here.

  6. parallelism here. Instead of these Instead of these sequential processes sequential processes , let's try to , let's try to , let's try to parallelize the work parallelize the work parallelize the work using a using a using a worker pool. So when you worker pool. So when you worker pool. So when you request 256 request 256 request 256 examples, it might be examples, it might be examples, it might be worth processing them all worth processing them all worth processing them all at once. Ahem. Yes. at once. Ahem. Yes. at once. Ahem. Yes. For example, we have two For example, we have two For example, we have two types of expenses. The first types of expenses. The first is obtaining is obtaining is obtaining images, the second is images, the second is images, the second is processing them. These are processing them. These are processing them. These are decoding, decoding, image preprocessing, and image preprocessing, and text tokenization. text tokenization. The samples are completely The samples are completely independent. This independent. This independent. This means that they are very means that they are very means that they are very easy to parallelize. easy to parallelize. So yes. Here you have to So yes. Here you have to make design make design make design decisions. What forms of decisions. What forms of decisions. What forms of parallelism do you parallelism do you parallelism do you want to use? There is want to use? There is want to use? There is asynchronous asynchronous asynchronous programming, which is programming, which is programming, which is suitable for suitable for suitable for input- input- output operations, such as SQL output operations, such as SQL queries or network queries or network queries or network requests. There are threads that requests. There are threads that requests. There are threads that can be can be can be used for used for CPU-intensive tasks and CPU-intensive tasks and free up the Python GIL, free up the Python GIL, free up the Python GIL, i.e. the global i.e. the global interpreter lock. It is not interpreter lock. It is not always obvious which always obvious which always obvious which operations operations operations release release release the lock. So you'll the lock. So you'll the lock. So you'll have to do have to do have to do a little more a little more a little more research. But research. But research. But usually it's usually it's usually it's using using using Hugging Face Hugging Face Hugging Face Transformers tokenizers, Transformers tokenizers, Transformers tokenizers, PIL transformation APIs, or PIL transformation APIs, or PIL transformation APIs, or some pandas operations.

  7. some pandas operations. Processes are more Processes are more resource-intensive, but resource-intensive, but resource-intensive, but provide complete provide complete provide complete isolation between isolation between isolation between workers. You can workers. You can workers. You can use use use multiple cores on your multiple cores on your multiple cores on your machine to get the machine to get the machine to get the job done. An in-depth job done. An in-depth job done. An in-depth analysis of the right analysis of the right analysis of the right solution deserves a solution deserves a solution deserves a separate report, separate report, separate report, but there are many but there are many but there are many online resources to online resources to online resources to help you help you help you decide on the decide on the decide on the choice for your choice for your choice for your case. What is important for case. What is important for case. What is important for our solution is our solution is our solution is the elimination of the “ the elimination of the “ the elimination of the “ bottleneck”. And not just bottleneck”. And not just bottleneck”. And not just using a using a using a tool that tool that tool that worked for us worked for us worked for us before. So, we will before. So, we will before. So, we will choose asynchronous I/ choose asynchronous I/ choose asynchronous I/ O O O for for for image retrieval, and image retrieval, and use use Ray actors for multiprocessing. Ray actors are essentially the Ray actors for multiprocessing. Ray actors are essentially the Ray actors for multiprocessing. Ray actors are essentially the same as regular same as regular same as regular processes, but they processes, but they processes, but they provide an API for provide an API for provide an API for interaction between interaction between interaction between workers, so we don't workers, so we don't workers, so we don't have to have to have to reinvent the wheel reinvent the wheel . So yes, your process is only as . So yes, your process is only as . So yes, your process is only as fast as fast as its slowest its slowest link. We can see that link. We can see that link. We can see that our wait time our wait time our wait time has decreased in the second has decreased in the second has decreased in the second example, which is example, which is example, which is highlighted in highlighted in highlighted in pink. You can pink. You can pink. You can see that the time to see that the time to see that the time to receive data for the receive data for the receive data for the packet has decreased to, well packet has decreased to, well , about 40 , about 40 , about 40 seconds. So, this gives seconds. So, this gives seconds. So, this gives us a 25% speedup.

  8. us a 25% speedup. However, as you can However, as you can see, there is a see, there is a see, there is a significant gap, significant gap, significant gap, especially when we especially when we especially when we go through many go through many go through many steps. So can steps. So can steps. So can we somehow limit the we somehow limit the we somehow limit the bottleneck bottleneck bottleneck associated with associated with associated with waiting for a request waiting for a request waiting for a request from the trainer before from the trainer before from the trainer before launching the launching the launching the worker pool to worker pool to worker pool to retrieve data? Maybe retrieve data? Maybe we should just we should just we should just get the data get the data get the data in advance? We in advance? We in advance? We proposed a second proposed a second proposed a second solution: the background solution: the background solution: the background thread pre- thread pre- thread pre- populates the queue and populates the queue and populates the queue and always stays ahead of always stays ahead of always stays ahead of the consumer. This means the consumer. This means that your producer that your producer that your producer creates examples creates examples creates examples faster than the consumer, faster than the consumer, faster than the consumer, so the trainer never so the trainer never so the trainer never waits for a data request. waits for a data request. When I need When I need data, it is available data, it is available data, it is available upon request. upon request. There is one There is one caveat to prior receipt. This makes it somewhat caveat to prior receipt. This makes it somewhat caveat to prior receipt. This makes it somewhat more difficult to more difficult to more difficult to create create create checkpoints and checkpoints and checkpoints and resume the resume the resume the model and model and model and data loader, data loader, data loader, and can also and can also and can also cause a "cold cause a "cold cause a "cold start" while the queue start" while the queue start" while the queue fills up for fills up for fills up for training. But training. But training. But considering these details, considering these details, considering these details, I think I think I think pre-ordering pre-ordering is a great idea.

  9. is a great idea. I think if your I think if your producer is always producer is always producer is always ahead of ahead of ahead of the consumer, it will pay the consumer, it will pay the consumer, it will pay you big dividends. It's you big dividends. It's no surprise that we no surprise that we achieved almost 0 achieved almost 0 achieved almost 0 seconds to receive seconds to receive seconds to receive and process data for and process data for and process data for each packet, each packet, each packet, as we had already as we had already as we had already received and processed received and processed received and processed it long before it long before the trainer needed it. Thus, the trainer needed it. Thus, we have effectively eliminated we have effectively eliminated we have effectively eliminated bottlenecks in both data bottlenecks in both data bottlenecks in both data loading and loading and loading and data transformation. data transformation. data transformation. So after these two So after these two So after these two steps, we decided we were steps, we decided we were steps, we decided we were done. done. done. Bandwidth Bandwidth Bandwidth has increased, GPUs are has increased, GPUs are has increased, GPUs are loaded. However, we loaded. However, we loaded. However, we noticed that noticed that noticed that data transfer costs still data transfer costs still data transfer costs still exist. In fact, exist. In fact, exist. In fact, they even grew. they even grew. they even grew. Each sample passes a Each sample passes a Each sample passes a large number of large number of large number of image arrays image arrays image arrays across process boundaries across process boundaries across process boundaries multiple times. Each multiple times. Each multiple times. Each sample is sample is pickled on the worker, pickled on the worker, pickled on the worker, copied to copied to copied to the coordinator, and then the coordinator, and then pickled again for pickled again for the trainer. For each the trainer. For each the trainer. For each sample, at each sample, at each sample, at each step, a step, a step, a memory copy occurs memory copy occurs memory copy occurs that no one really that no one really that no one really needed. This needed. This needed. This causes causes causes your your your data generator to frequently throw an out of data generator to frequently throw an out of memory (OOM) error.

  10. Multimodal Multimodal tensors are huge. tensors are huge. tensors are huge. They take up They take up They take up megabytes per megabytes per megabytes per sample. By sample. By sample. By default, each default, each default, each sample is returned sample is returned sample is returned by value, which by value, which by value, which means that these heavy means that these heavy means that these heavy arrays are copied arrays are copied arrays are copied between processes. So, between processes. So, object storage could be the solution for this object storage could be the solution for this . This is a convenient . This is a convenient . This is a convenient method that Ray provides you with method that Ray provides you with method that Ray provides you with . Each sample . Each sample . Each sample contains a reference to contains a reference to contains a reference to an object, which is simply a an object, which is simply a an object, which is simply a pointer to data. pointer to data. pointer to data. The driver simply The driver simply The driver simply stores this stores this stores this pointer. The tensor pointer. The tensor pointer. The tensor never never never returns to the returns to the returns to the master node. master node. Megatron's data parallelism ranks can Megatron's data parallelism ranks can directly directly directly access access access Ray storage to Ray storage to Ray storage to retrieve the required retrieve the required retrieve the required data. So, whenever data. So, whenever data. So, whenever choosing a choosing a choosing a solution for solution for solution for parallelism and parallelism and parallelism and concurrency, you concurrency, you concurrency, you should always should always should always consider the consider the consider the additional additional additional communication costs. So communication costs. So communication costs. So plan it wisely plan it wisely . OK. Yes, we . OK. Yes, we . OK. Yes, we were able to eliminate the were able to eliminate the were able to eliminate the redundant " redundant " put object" calls that were in put object" calls that were in put object" calls that were in our prefetch solution our prefetch solution by by using using using Ray storage. We see that our Ray storage. We see that our Ray storage. We see that our communication communication communication costs have decreased and the costs have decreased and the costs have decreased and the proportion of proportion of proportion of waiting time has also waiting time has also waiting time has also decreased. So, decreased. So, decreased. So, to summarize: we took a to summarize: we took a to summarize: we took a sequential framework, sequential framework, sequential framework, added parallelism added parallelism added parallelism for loading and for loading and for loading and processing data, added processing data, added processing data, added prefetching prefetching prefetching to get ahead of to get ahead of to get ahead of training requests, training requests, training requests, and added Ray storage and added Ray storage and added Ray storage to minimize to minimize to minimize data transfer. We data transfer. We data transfer. We have reduced the base have reduced the base expectation ratio from 85% to 20%.

  11. You can see that the You can see that the sawtooth graph sawtooth graph sawtooth graph above has largely above has largely above has largely flattened out, which flattened out, which flattened out, which means that means that means that our GPU usage our GPU usage our GPU usage has increased. So as we has increased. So as we has increased. So as we scale, scale, scale, now that we have an efficient now that we have an efficient multimodal multimodal learning pipeline, let's learning pipeline, let's learning pipeline, let's try to try to try to scale it to a scale it to a scale it to a production run. During production run. During production run. During this, we noticed this, we noticed that the wait time fraction that the wait time fraction that the wait time fraction for the same for the same for the same pipeline increased, even though pipeline increased, even though pipeline increased, even though we just changed we just changed we just changed some numbers. We some numbers. We some numbers. We quadrupled the quadrupled the quadrupled the ranks of ranks of ranks of data parallelism, added about data parallelism, added about data parallelism, added about 200 data streams, and 200 data streams, and quadrupled the batch size quadrupled the batch size . It turns out that . It turns out that . It turns out that because the because the because the data generator and the worker pool data generator and the worker pool data generator and the worker pool are located together are located together are located together for fast for fast for fast data transfer, a bottleneck occurs data transfer, a bottleneck occurs data transfer, a bottleneck occurs on the network on the network on the network map of this node. This map of this node. This map of this node. This means that when we means that when we means that when we concentrate all this concentrate all this concentrate all this work and processing on work and processing on work and processing on certain nodes, their certain nodes, their network bandwidth cannot keep up with the network bandwidth cannot keep up with the needs of the needs of the needs of the learning pipeline. At this learning pipeline. At this learning pipeline. At this point, we point, we point, we dug deeper into the frameworks we dug deeper into the frameworks we were using were using were using and saw that by and saw that by and saw that by default Ray default Ray default Ray schedules workers schedules workers schedules workers mostly on the mostly on the mostly on the same nodes, which same nodes, which same nodes, which leads to leads to leads to easy easy bottlenecks on bottlenecks on bottlenecks on one machine before one machine before one machine before moving to moving to moving to another. So, we another. So, we another. So, we looked at looked at looked at the possibility the possibility the possibility of disaggregating of disaggregating of disaggregating this entire process, where this entire process, where this entire process, where worker nodes

  12. worker nodes worker nodes are distributed across are distributed across are distributed across multiple servers multiple servers multiple servers so that they can so that they can so that they can independently process independently process independently process and store and store and store images, which are then images, which are then retrieved and retrieved and used by used by used by our DPU rings in a timely manner. our DPU rings in a timely manner. So yes, these are two So yes, these are two parameters, two parameters, two parameters, two checkboxes that we checkboxes that we checkboxes that we actually set. actually set. actually set. The first is The first is The first is zero- zero- zero- copy data retrieval, and the second is a copy data retrieval, and the second is a distributed distributed scheduling mechanism. None of scheduling mechanism. None of scheduling mechanism. None of them helped in a them helped in a them helped in a small-scale small-scale small-scale experiment, but experiment, but experiment, but together they produced a together they produced a together they produced a synergistic effect synergistic effect synergistic effect in a large-scale in a large-scale in a large-scale experiment. experiment. A configuration that A configuration that proves proves proves ineffective on a small ineffective on a small ineffective on a small scale may scale may scale may work very well work very well work very well on a large scale. Therefore, it is on a large scale. Therefore, it is on a large scale. Therefore, it is always worth thinking always worth thinking always worth thinking about how to conduct about how to conduct about how to conduct ablation ablation ablation studies at studies at studies at different levels different levels different levels of scale. It was of scale. It was of scale. It was these two flags that these two flags that these two flags that ultimately significantly ultimately significantly ultimately significantly affected affected affected the speed of the speed of the speed of our our our data processing pipeline. This is the data processing pipeline. This is the data processing pipeline. This is the final version of final version of final version of the pipeline that we have the pipeline that we have the pipeline that we have reached. An up-to-date, reached. An up-to-date, reached. An up-to-date, operational system operational system operational system that takes into account that takes into account that takes into account data access patterns, data access patterns, data access patterns, distributes training distributes training distributes training between workloads, between workloads, between workloads, and keeps the GPU and keeps the GPU and keeps the GPU active, active, active, ensuring ensuring ensuring smooth and smooth and smooth and stable stable stable training. Overall, we training. Overall, we training. Overall, we got 50% more got 50% more got 50% more bandwidth bandwidth bandwidth at scale just from at scale just from at scale just from these two these two these two flags. This is the flags. This is the flags. This is the same lever, but with the same lever, but with the same lever, but with the opposite opposite opposite result, result, result, defined by the defined by the defined by the bottleneck. Here are the main bottleneck. Here are the main bottleneck. Here are the main conclusions that can be conclusions that can be conclusions that can be drawn from this drawn from this drawn from this report. GPUs are report. GPUs are report. GPUs are expensive, but sometimes the

  13. expensive, but sometimes the expensive, but sometimes the most important battles most important battles most important battles take place outside of the take place outside of the take place outside of the GPU. So to GPU. So to GPU. So to start, ask start, ask start, ask yourself: what is the yourself: what is the yourself: what is the bottleneck—CPU or GPU? Are you bottleneck—CPU or GPU? Are you bottleneck—CPU or GPU? Are you limited by limited by limited by I/O I/O I/O or computing or computing or computing power? power? power? Moreover, solving one Moreover, solving one Moreover, solving one bottleneck often bottleneck often bottleneck often opens up another. opens up another. opens up another. First for us it First for us it First for us it was was was bandwidth, then bandwidth, then bandwidth, then latency, then latency, then latency, then memory, and then memory, and then memory, and then the cost of getting the the cost of getting the the cost of getting the data. Classical data. Classical data. Classical approaches are not everything; approaches are not everything; approaches are not everything; Parallelism and Parallelism and preloading are important , and many have , and many have , and many have already talked about this, but it is equally already talked about this, but it is equally already talked about this, but it is equally important important important to consider the to consider the to consider the semantics semantics semantics of the frameworks and the of the frameworks and the of the frameworks and the constraints you are constraints you are constraints you are working with. Scale working with. Scale working with. Scale always changes always changes always changes the answer. the answer. the answer. Setting Setting Setting parameters such as parameters such as parameters such as zero-copy, zero-copy, zero-copy, wait, and wait, and wait, and worker node distribution worker node distribution only become important at only become important at large scales and large scales and large scales and may not yield may not yield may not yield noticeable improvements noticeable improvements noticeable improvements at lower levels. at lower levels. Overall, Overall, baseline profiling baseline profiling baseline profiling helps helps helps achieve shifts. achieve shifts. achieve shifts. Thank you all for Thank you all for Thank you all for attending this attending this attending this talk. If you have any talk. If you have any talk. If you have any questions, I will be questions, I will be questions, I will be happy to happy to happy to answer them. Thank you.

No summary available yet.

View original episode ↗