This is a commercial-grade digital human video application based on Wan + InfiniteTalk.
After users upload a character image and audio, the workflow first performs vocal separation and audio encoding on the audio, then combines it with the character image, prompt, size control, and multi-round sampling to generate a digital human video with lip-sync, natural expressions, and higher coherence. The current file clearly shows that it includes LoadAudio, LoadImage, MelBandRoFormerSampler, AudioEncoderEncode, WanInfiniteTalkToVideo, multiple groups of KSamplerAdvanced, VHS_VideoCombine, as well as groups like Character1 Audio and Mask, Character2 Audio and Mask, 分辨率 (Resolution), and 每次处理帧数 (Frames per batch). This indicates that it is not a simple single-stage lip-sync driver, but a complete workflow biased toward commercial finished video production.
This workflow is suitable for:
Recommended usage:
Upload a clear character image with a relatively complete face and avoid extreme angles; upload clear audio; prompts should ideally describe the character's speaking state, emotion, and camera atmosphere, while avoiding overly exaggerated actions. Since this workflow already includes audio separation, audio encoding, character image size adaptation, segmented loop generation, and final audio-video synthesis, it is better suited for applications where "users only provide the image and audio to directly produce a finished digital human video."
🎁Claim RH Coins first, then experience the workflow! Click the avatar in the top right corner → Invitation Code → Enter [rh-v1111] to instantly receive 1000 RH Coins. Log in daily to claim another 100 coins~ For more ComfyUI workflows, tutorials, and gameplay, please follow our Official WeChat Account (AIKSK), with simultaneous updates on Douyin / Bilibili / Xiaohongshu / YouTube (AI-KSK) ✨
No creations yet

No creations available.