
ltx-2.3-22b-dev_transformer_only_mxfp8_block32.safetensors
Model Information
This can be considered the current definitive solution that balances "lightning-fast generation" and "ultimate image quality (under low VRAM)."
This mxfp8_block32 model does not use traditional global scaling; instead, it divides the weights into blocks of 32 elements each for micro-scaling.
Advantages: This block-wise quantization greatly preserves the original distribution characteristics of the model. While maintaining almost the exact same file size and VRAM usage as standard FP8, its image quality approaches the full-blood BF16 infinitely. It effectively resolves the flickering and noise issues commonly seen in standard FP8 quantization during video generation.
transformer_only means that it only contains the core backbone network of LTX 2.3 (UNet/DiT), keeping VRAM usage low while allowing users to optionally pair it with higher-precision text_projection_bf16.
---
This can be considered the current solution that balances "extremely fast image output" and "ultimate image quality (with low VRAM)".
This mxfp8_block32 model does not use traditional global scaling. Instead, it divides the weights into blocks of 32 elements each for micro-scaling.
Advantages: This block quantization greatly preserves the original distribution characteristics of the model. While having almost the same file size and VRAM usage as ordinary FP8, its image quality is nearly indistinguishable from full-fledged BF16. It effectively solves the flickering and noise problems commonly found in ordinary FP8 quantization during video generation.
"Transformer_only" means that they only include the core backbone network of LTX 2.3 (UNet/DiT), allowing for low VRAM usage while being compatible with higher-precision text_projection_bf16.
This can be considered the current definitive solution that balances "lightning-fast generation" and "ultimate image quality (under low VRAM)."
This mxfp8_block32 model does not use traditional global scaling; instead, it divides the weights into blocks of 32 elements each for micro-scaling.
Advantages: This block-wise quantization greatly preserves the original distribution characteristics of the model. While maintaining almost the exact same file size and VRAM usage as standard FP8, its image quality approaches the full-blood BF16 infinitely. It effectively resolves the flickering and noise issues commonly seen in standard FP8 quantization during video generation.
transformer_only means that it only contains the core backbone network of LTX 2.3 (UNet/DiT), keeping VRAM usage low while allowing users to optionally pair it with higher-precision text_projection_bf16.
---
This can be considered the current solution that balances "extremely fast image output" and "ultimate image quality (with low VRAM)".
This mxfp8_block32 model does not use traditional global scaling. Instead, it divides the weights into blocks of 32 elements each for micro-scaling.
Advantages: This block quantization greatly preserves the original distribution characteristics of the model. While having almost the same file size and VRAM usage as ordinary FP8, its image quality is nearly indistinguishable from full-fledged BF16. It effectively solves the flickering and noise problems commonly found in ordinary FP8 quantization during video generation.
"Transformer_only" means that they only include the core backbone network of LTX 2.3 (UNet/DiT), allowing for low VRAM usage while being compatible with higher-precision text_projection_bf16.