This is a single-image audio-driven singing video application based on WanVideo + InfiniteTalk / MultiTalk.
After a user uploads a character image and an audio clip, the workflow first standardizes the image dimensions, uses a large model to organize prompt terms based on the visual content and user input, then encodes the text using WanVideoTextEncodeSingle, extracts audio-driven features using MultiTalkWav2VecEmbeds, and finally passes everything to the WanVideo base model to generate a video where the character sings along with the audio, featuring subtle facial expressions and movement changes. The entire pipeline simultaneously incorporates structures like WanVideoModelLoader, WanVideoTextEncodeSingle, MultiTalkWav2VecEmbeds, WanVideoApplyNAG, and WanVideoVAELoader, indicating that this is not a standard image-to-video workflow, but rather one tailored for audio-driven digital human singing/rapping.
This workflow is suitable for:
Recommended usage:
Upload a clear front-facing portrait with a well-defined subject; upload a clear vocal audio file; keep the prompts focused on "how the character sings/speaks in what setting" rather than overly exaggerated big movements. The current file also connects to a llama_cpp_instruct_adv pipeline, which first processes the image and prompts before feeding them into the final generation flow, giving this workflow the built-in capability to "help users organize prompts."
🎁Claim RH Coins first, then experience the workflow! Click the avatar in the top right corner → Invitation Code → Enter [rh-v1111] to instantly receive 1000 RH Coins, and log in daily to claim another 100 RH Coins~ For more ComfyUI workflows, tutorials, and gameplay, please follow our official WeChat account (AIKSK), with synchronized updates on TikTok/Douyin / Bilibili / Xiaohongshu / YouTube (AI-KSK) ✨
No creations yet

No creations available.