The FantasyTalking project, jointly launched by Alibaba and Beijing University of Posts and Telecommunications, marks another major breakthrough in digital human technology. With just one ID photo, it can generate digital human videos with expressive emotions and natural movements.
Three Innovative Modules
- Audio-Visual Alignment Strategy: Captures the global correlation between audio, facial expressions, body movements, and background dynamics.
- Facial Cross-Attention: Locks in facial features with only 3% of parameter volume, achieving an identity shift rate of <0.3% for 10-minute videos.
- Motion Intensity Modulation Network: Independently controls 22 parameters for facial/body movement amplitudes (e.g., eyebrow height, shoulder swing frequency).
Breakthrough in Generation Effects
- Supports 9 Generation Modes: Close-up/half-body/full-body, frontal/side view, dynamic background.
- Covers multiple styles including realistic humans/cartoons/animals, with lip-sync error <40ms.
- 360° surround view generation, with realistic details such as hair fluttering and neck folds.
Performance Comparison Advantages
In the OmniHuman 1 benchmark test, it leads in motion coherence (CIDEr↑18%) and identity preservation (SSIM↑23%).
Model download link: https://pan.quark.cn/s/184684a6d030