
ltx-2.3-22b-dev-fp8-Lightricks-官方版
Model Information
ltx-2.3-22b-dev-fp8-Lightricks-Official Version
-----------------------------------------------
LTX-2.3 Model Series / Full Download
Files shared via cloud drive: LTX2.3 Model Full Set
Link: https://pan.baidu.com/s/1hN4rDcNqhAKb1_PLsnIGyg?pwd=77ra Extraction Code: 77ra
-----------------------------------------------
LTX-2.3 is a new generation of open-source audio-video integrated generation models developed by Lightricks. As an upgraded version of the LTX-2 series, it has been launched and released on the Alibaba Cloud ModelScope community, along with an FP8 quantized and optimized version (ltx-2.3-fp8). It is primarily designed for AIGC video creation scenarios such as Text-to-Video (T2V) and Image-to-Video (I2V), making it a multimodal generation model that balances generation quality, engineering controllability, and deployment practicality.
Core R&D Background and Positioning
Continuing the core architecture of the LTX series based on Diffusion Transformers (DiT), LTX-2.3 is a comprehensive functional and quality upgrade to LTX-2. It focuses on the core capability of audio-video synchronized generation. Unlike traditional single-video generation models, it enables integrated video + audio creation under text/image inputs. At the same time, it optimizes dynamic performance, image quality, and compatibility based on the practical needs of the open-source community. It is positioned as an open-source video generation solution for industrial-grade and individual creators, balancing free usage, engineering controllability, and localized deployment features.
The model provides a base version, a quantized version (FP8), and supporting LoRA/IC-LoRA models, adapting to different hardware resources and use cases. It currently supports mainstream open-source frameworks such as ComfyUI and Hugging Face Diffusers, offering strong out-of-the-box usability.
Core Upgrades and Feature Highlights
As an iterative version of LTX-2, LTX-2.3 achieves multi-dimensional improvements in generation quality, feature support, and detail performance, which are the core competitiveness of this version:
Completely resolves audio clipping issues: Optimizations have been made regarding the audio distortion and clipping issues in previous versions during text/video generation. Audio-visual synchronization has been significantly improved, and the naturalness of character speech and scene sound effects is markedly enhanced.
Upgraded dynamic performance and camera effects: Camera movement and cutting logic are more rational. The fluidity of large-scale dynamics (such as boxing, racing, and cyberpunk cityscapes) is greatly enhanced, and motion continuity performs excellently among open-source video generation models.
Native support for vertical screen scenarios: Added adaptation for vertical video generation, catering to the mainstream demands of short video and social media creation, allowing the generation of video content suited for mobile displays without extra adjustments.
Improved character facial generation stability: Even when characters occupy a relatively small proportion of the frame, it effectively avoids facial degradation and distortion, resulting in more natural generation for realistic character scenes.
Chinese scenario adaptation optimization: Achieves automatic subtitle generation for Chinese videos for the first time (English scenarios are not yet supported). Although subtitle recognition still has minor flaws, it completes core adaptation for Chinese creation scenarios.
Image quality and detail optimization: Image quality is excellent in static/low-dynamic scenes such as landscapes, realism, and horror styles, with high detail restoration for 10-second short videos.
Model Versions and Formats
Lightricks provides multi-format model packages for LTX-2.3 to suit different usage requirements, all updated in the ModelScope community:
Base Version (LTX-2.3): Full-featured version containing 3 base models and supporting LoRA/IC-LoRA. The main model file is approximately 40+ GB, retaining all generation capabilities for optimal results.
FP8 Quantized Version (ltx-2.3-fp8): A low-precision version optimized for VRAM usage and inference speed, drastically lowering the hardware threshold, suitable for resource-constrained localized deployment.
Applicable Scenarios
LTX-2.3 inherits the multi-scenario adaptability of the LTX series. Combined with its upgraded features, it better fits practical demands such as short video creation, commercial content production, and personalized creative generation. Core applicable scenarios include:
Social Media / Self-Media Creation: Rapidly generate vertical short videos, narrative shorts, and landscape/creative videos, supporting automatic Chinese subtitle generation to fit platform needs like Douyin and Xiaohongshu.
Commercial Advertising and Product Display: Quickly generate dynamic e-commerce product display videos and brand advertising shorts via images/text, reducing production costs.
Education and Training: Teachers can generate dynamic teaching videos using text prompts, combined with synchronized audio explanations to enrich teaching formats.
Gaming and Virtual Content: Generate dynamic animations and matching sound effects for game characters and virtual scenes, enhancing immersion in virtual worlds.
Artistic Creation and Visual Storytelling: Supports multiple styles such as cyberpunk, realism, and horror, meeting creators' personalized artistic expression needs.
Current Limitations
LTX-2.3 is still an optimized version in the open-source stage with a few imperfect issues to keep in mind during use:
Poor anime/2D generation results: While realistic scenes perform excellently, video generation for 2D/anime styles is subpar. Officials recommend using the Image-to-Video (I2V) mode rather than Text-to-Video (T2V) for anime creation.
High-dynamic scene quality degradation: High-motion frames (such as high-speed movement and complex scene transitions) may exhibit slight "graininess / pixelation" issues, with quality slightly lower than low-dynamic scenes.
LoRA parameters require fine-tuning: When using the accompanying LoRA at maximum intensity (set to 1), character faces are prone to aging and distortion, requiring reduced intensity alongside sampler adjustments.
Limited subtitle generation accuracy: While Chinese automatic subtitles can be generated, some text recognition errors and incomplete parts exist, requiring post-processing proofreading.
Video duration currently limited to short frames: Optimal generation results are concentrated around 10-second short videos, and scene consistency for longer videos still has room for improvement.
Deployment and Usage Adaptation
Framework Support: Natively supports ComfyUI (graphical operation, suitable for individual creators) and Hugging Face Diffusers (underlying code library, suitable for developers' programmatic deployment). Workflows can be quickly set up via one-click integration packages.
Hardware Requirements: The base version main model is about 40+ GB; high-performance GPUs (such as NVIDIA RTX 4090/50 series, A100/H100) are recommended. The FP8 quantized version lowers the hardware threshold, supporting local execution on mid-to-low-end config GPUs.
Tips for Use: Lowering distillation LoRA intensity during generation, using a standard sampler for 4 steps, and setting the denoise value to 0.3-0.5 can effectively optimize image quality and character facial performance.
Core Advantages Over Similar Open-Source Models
This model is sourced from an external transfer (transfer address: https://www.modelscope.cn/models/Lightricks/ ),if the original author has objections to this transfer, you can click,
ltx-2.3-22b-dev-fp8-Lightricks-Official Version
-----------------------------------------------
LTX-2.3 Model Series / Full Download
Files shared via cloud drive: LTX2.3 Model Full Set
Link: https://pan.baidu.com/s/1hN4rDcNqhAKb1_PLsnIGyg?pwd=77ra Extraction Code: 77ra
-----------------------------------------------
LTX-2.3 is a new generation of open-source audio-video integrated generation models developed by Lightricks. As an upgraded version of the LTX-2 series, it has been launched and released on the Alibaba Cloud ModelScope community, along with an FP8 quantized and optimized version (ltx-2.3-fp8). It is primarily designed for AIGC video creation scenarios such as Text-to-Video (T2V) and Image-to-Video (I2V), making it a multimodal generation model that balances generation quality, engineering controllability, and deployment practicality.
Core R&D Background and Positioning
Continuing the core architecture of the LTX series based on Diffusion Transformers (DiT), LTX-2.3 is a comprehensive functional and quality upgrade to LTX-2. It focuses on the core capability of audio-video synchronized generation. Unlike traditional single-video generation models, it enables integrated video + audio creation under text/image inputs. At the same time, it optimizes dynamic performance, image quality, and compatibility based on the practical needs of the open-source community. It is positioned as an open-source video generation solution for industrial-grade and individual creators, balancing free usage, engineering controllability, and localized deployment features.
The model provides a base version, a quantized version (FP8), and supporting LoRA/IC-LoRA models, adapting to different hardware resources and use cases. It currently supports mainstream open-source frameworks such as ComfyUI and Hugging Face Diffusers, offering strong out-of-the-box usability.
Core Upgrades and Feature Highlights
As an iterative version of LTX-2, LTX-2.3 achieves multi-dimensional improvements in generation quality, feature support, and detail performance, which are the core competitiveness of this version:
Completely resolves audio clipping issues: Optimizations have been made regarding the audio distortion and clipping issues in previous versions during text/video generation. Audio-visual synchronization has been significantly improved, and the naturalness of character speech and scene sound effects is markedly enhanced.
Upgraded dynamic performance and camera effects: Camera movement and cutting logic are more rational. The fluidity of large-scale dynamics (such as boxing, racing, and cyberpunk cityscapes) is greatly enhanced, and motion continuity performs excellently among open-source video generation models.
Native support for vertical screen scenarios: Added adaptation for vertical video generation, catering to the mainstream demands of short video and social media creation, allowing the generation of video content suited for mobile displays without extra adjustments.
Improved character facial generation stability: Even when characters occupy a relatively small proportion of the frame, it effectively avoids facial degradation and distortion, resulting in more natural generation for realistic character scenes.
Chinese scenario adaptation optimization: Achieves automatic subtitle generation for Chinese videos for the first time (English scenarios are not yet supported). Although subtitle recognition still has minor flaws, it completes core adaptation for Chinese creation scenarios.
Image quality and detail optimization: Image quality is excellent in static/low-dynamic scenes such as landscapes, realism, and horror styles, with high detail restoration for 10-second short videos.
Model Versions and Formats
Lightricks provides multi-format model packages for LTX-2.3 to suit different usage requirements, all updated in the ModelScope community:
Base Version (LTX-2.3): Full-featured version containing 3 base models and supporting LoRA/IC-LoRA. The main model file is approximately 40+ GB, retaining all generation capabilities for optimal results.
FP8 Quantized Version (ltx-2.3-fp8): A low-precision version optimized for VRAM usage and inference speed, drastically lowering the hardware threshold, suitable for resource-constrained localized deployment.
Applicable Scenarios
LTX-2.3 inherits the multi-scenario adaptability of the LTX series. Combined with its upgraded features, it better fits practical demands such as short video creation, commercial content production, and personalized creative generation. Core applicable scenarios include:
Social Media / Self-Media Creation: Rapidly generate vertical short videos, narrative shorts, and landscape/creative videos, supporting automatic Chinese subtitle generation to fit platform needs like Douyin and Xiaohongshu.
Commercial Advertising and Product Display: Quickly generate dynamic e-commerce product display videos and brand advertising shorts via images/text, reducing production costs.
Education and Training: Teachers can generate dynamic teaching videos using text prompts, combined with synchronized audio explanations to enrich teaching formats.
Gaming and Virtual Content: Generate dynamic animations and matching sound effects for game characters and virtual scenes, enhancing immersion in virtual worlds.
Artistic Creation and Visual Storytelling: Supports multiple styles such as cyberpunk, realism, and horror, meeting creators' personalized artistic expression needs.
Current Limitations
LTX-2.3 is still an optimized version in the open-source stage with a few imperfect issues to keep in mind during use:
Poor anime/2D generation results: While realistic scenes perform excellently, video generation for 2D/anime styles is subpar. Officials recommend using the Image-to-Video (I2V) mode rather than Text-to-Video (T2V) for anime creation.
High-dynamic scene quality degradation: High-motion frames (such as high-speed movement and complex scene transitions) may exhibit slight "graininess / pixelation" issues, with quality slightly lower than low-dynamic scenes.
LoRA parameters require fine-tuning: When using the accompanying LoRA at maximum intensity (set to 1), character faces are prone to aging and distortion, requiring reduced intensity alongside sampler adjustments.
Limited subtitle generation accuracy: While Chinese automatic subtitles can be generated, some text recognition errors and incomplete parts exist, requiring post-processing proofreading.
Video duration currently limited to short frames: Optimal generation results are concentrated around 10-second short videos, and scene consistency for longer videos still has room for improvement.
Deployment and Usage Adaptation
Framework Support: Natively supports ComfyUI (graphical operation, suitable for individual creators) and Hugging Face Diffusers (underlying code library, suitable for developers' programmatic deployment). Workflows can be quickly set up via one-click integration packages.
Hardware Requirements: The base version main model is about 40+ GB; high-performance GPUs (such as NVIDIA RTX 4090/50 series, A100/H100) are recommended. The FP8 quantized version lowers the hardware threshold, supporting local execution on mid-to-low-end config GPUs.
Tips for Use: Lowering distillation LoRA intensity during generation, using a standard sampler for 4 steps, and setting the denoise value to 0.3-0.5 can effectively optimize image quality and character facial performance.
Core Advantages Over Similar Open-Source Models
This model is sourced from an external transfer (transfer address: https://www.modelscope.cn/models/Lightricks/ ),if the original author has objections to this transfer, you can click,