This is a multi-element audio-driven video application based on LTX 2.3 Dev-Dare + 4-Image Reference Guidance + Driving Audio + PromptRelay Timeline Control + NAG Audio-Visual Condition Enhancement + Three-Stage Sampling Refinement. After users upload 4 reference images and 1 driving audio track, the system injects the four images as the background, character, props, and interactive elements respectively into the video generation pipeline, and then generates a dynamic video with audio-driven rhythm combined with the audio beats.
The core advantages of this workflow are: the four images are responsible for visual elements, the audio is responsible for rhythm and speaking feel, and PromptRelay is responsible for timeline actions. Compared to standard 4-image reference videos, it features additional audio latent fusion and audio-visual condition enhancement, making it ideal for character lip-syncing, character-prop interaction, character-pet/sprite interaction, product voiceover showcases, AI short drama dialogues, virtual human marketing, music rhythm videos, and multi-element narrative shots. When using it, it is recommended to focus on controlling four aspects: the 4 reference images determine the element sources, the driving audio determines rhythm and speaking feel, the global prompt (used to write overall visual rules) determines element relationships, and the segmented prompts (writing action changes for each time period) determine the interaction rhythm of each segment.
🎁Claim RH Coins first, then try out the workflow! Go to Avatar in the top right corner → Invitation Code → Enter [rh-v1111] to instantly get 1000 RH Coins. Log in daily to claim another 100 RH Coins~ For more ComfyUI workflows, tutorials, and gameplay, please follow our Official Account (AIKSK), with simultaneous updates on Douyin / Bilibili / Xiaohongshu / YouTube (AI-KSK)✨


No creations yet

No creations available.