Wan2.1 is the next-generation video generation model developed by Alibaba Tongyi Wanxiang team, achieving significant breakthroughs in AI-driven visual content creation.
- • Bilingual video model: Wan2.1 is the first video model capable of generating bilingual (Chinese and English) text, with powerful text generation capabilities that enhance its practicality. It can create text and animations with cinematic-level effects, supporting font applications in various scenarios, including special effect fonts, poster fonts, and font displays in real-world scenes, meeting diverse professional needs.
- • Multi-video tasks: Provides robust capabilities for text-to-video and image-to-video generation, as well as video editing, video-to-audio, and other tasks.
- • High-quality performance: Wan2.1 is based on a hybrid Variational Autoencoder (VAE) and Diffusion Transformer (DiT) architecture, enhancing temporal modeling and scene understanding capabilities. Through multimodal fusion technology, it can simultaneously generate high-definition videos, dynamic subtitles, and multilingual dubbing, supporting 1080p resolution and efficient encoding and decoding to ensure high-quality video output. In January 2025, Alibaba Tongyi Wanxiang Wan2.1 model topped the Vbench leaderboard, surpassing Sora, HunyuanVideo, Minimax, Luma, Gen3, Pika, and other domestic and international video generation models. It also continuously outperforms existing open-source models and state-of-the-art commercial solutions in multiple benchmark tests.