← Back
Aaron Zisk June 19, 2026 1m

This Is What Happens When You CRUSH An AI Video Model

Read full transcript 2 segments
  1. Here's a thing that nobody tells you. If Here's a thing that nobody tells you. If you're running an AI video model you're running an AI video model you're running an AI video model locally, you're running it quantized. I locally, you're running it quantized. I locally, you're running it quantized. I took two video models and I ran each one took two video models and I ran each one took two video models and I ran each one all the way down an eight-step ladder. all the way down an eight-step ladder. all the way down an eight-step ladder. When 2.2 14 billion parameters text to When 2.2 14 billion parameters text to When 2.2 14 billion parameters text to video and LTX 2.3 22 billion parameter video and LTX 2.3 22 billion parameter video and LTX 2.3 22 billion parameter model. Now we cut to half the bits and model. Now we cut to half the bits and model. Now we cut to half the bits and this is where it gets weird. Eight bits this is where it gets weird. Eight bits this is where it gets weird. Eight bits per weight but specifically the FP8 per weight but specifically the FP8 per weight but specifically the FP8 format. So floating points, half the format. So floating points, half the format. So floating points, half the bytes of FP16. The FP16 version looks bytes of FP16. The FP16 version looks bytes of FP16. The FP16 version looks just a little bit more realistic to me just a little bit more realistic to me just a little bit more realistic to me except for what are these lines walking except for what are these lines walking except for what are these lines walking across the table? I don't know, it might across the table? I don't know, it might across the table? I don't know, it might be raining or something or something be raining or something or something be raining or something or something else is moving in the room. On the else is moving in the room. On the else is moving in the room. On the right, it's a little choppier and a right, it's a little choppier and a right, it's a little choppier and a little bit more blurry. But if it wasn't little bit more blurry. But if it wasn't little bit more blurry. But if it wasn't next to the FP16 image, it would just next to the FP16 image, it would just next to the FP16 image, it would just pass. Here's how things change. FP8 is pass. Here's how things change. FP8 is pass. Here's how things change. FP8 is already off the baseline here. This is a already off the baseline here. This is a already off the baseline here. This is a static detail video. 0.19 on LPIPS. static detail video. 0.19 on LPIPS. static detail video. 0.19 on LPIPS. Here's the surprise. The very next level Here's the surprise. The very next level Here's the surprise. The very next level down, it's not really down, it's down, it's not really down, it's down, it's not really down, it's parallel I'd say, it's Q80. Q8 pretty parallel I'd say, it's Q80. Q8 pretty parallel I'd say, it's Q80. Q8 pretty much the same disk size as FP8 but much the same disk size as FP8 but much the same disk size as FP8 but better fidelity. And guess what? The better fidelity. And guess what? The better fidelity. And guess what? The average across all the five prompts average across all the five prompts average across all the five prompts here, FP8 is still about twice as far here, FP8 is still about twice as far here, FP8 is still about twice as far from baseline as Q8. So FP8, the format from baseline as Q8. So FP8, the format from baseline as Q8. So FP8, the format with hardware acceleration that with hardware acceleration that with hardware acceleration that everybody said was the future, it drifts everybody said was the future, it drifts everybody said was the future, it drifts further and further from full precision.

  2. further and further from full precision. further and further from full precision. Same number of bits, the difference is Same number of bits, the difference is Same number of bits, the difference is just where it spends that precision. And just where it spends that precision. And just where it spends that precision. And here is the kicker. The red car, a here is the kicker. The red car, a here is the kicker. The red car, a simple motion prompt. At FP8 the car simple motion prompt. At FP8 the car simple motion prompt. At FP8 the car drives backwards. Not in any other drives backwards. Not in any other drives backwards. Not in any other quantization does this happen. Just FP8. quantization does this happen. Just FP8. quantization does this happen. Just FP8. And this time, it's not just slightly And this time, it's not just slightly And this time, it's not just slightly different rendering, it's just wrong. So different rendering, it's just wrong. So different rendering, it's just wrong. So more bits is not the same as better. The more bits is not the same as better. The more bits is not the same as better. The format does more work than the bit count format does more work than the bit count format does more work than the bit count does. And if you remember one thing from does. And if you remember one thing from does. And if you remember one thing from this whole video, let it be this. Format this whole video, let it be this. Format this whole video, let it be this. Format matters more than bit count.

Summary

The main theme is the surprising impact of quantization formats on AI video model performance, particularly the FP8 format. Key subjects include specific video models, bit counts, and the FP8 and Q8 formats, with the unexpected finding that FP8 can introduce significant rendering errors. The practical takeaway is that format choice is more crucial for model fidelity than the raw bit count, emphasizing that "more bits is not the same as better."

View original episode ↗