This is a multi-image scene-cutting digital human application based on LongCat-Avatar-15 + LongCat DMD Distill LoRA + Whisper Large V3 audio lip-sync features + 10-image reference pool + 10 sets of GPT-5 image-to-prompt pools + For-loop continuation + overlapping frame stitching + strong scene-cut mode. After a user uploads an audio clip and prepares multiple digital human reference images, the system automatically converts the audio into lip-sync features, and then uses LongCat Avatar to generate the talking-head video segment by segment. Each cycle can switch reference images and corresponding prompts, thereby achieving multi-shot, multi-persona, and multi-scene talking-head videos.
The core advantage of this workflow is: It is not just a single image from start to finish, but a digital human talking-head production line featuring "audio-driven + multi-image rotation + loop continuation + automatic stitching". The first segment uses image 1 and the initial prompt to establish the character and visual, and subsequent segments continue generation via a For loop. When the reference-image mode is set to 1, images are rotated from the 10-image pool each round; when the prompt mode is set to 1, corresponding descriptions are rotated from the 10 sets of GPT-5 reverse-prompt pools each round; when the strong-cut mode is set to 1, it cuts more noticeably into the current reference image, making it ideal for multi-shot switching effects. It is particularly suitable for virtual anchors, multi-shot talking heads, AI news broadcasting, course narration, product introductions, character-rotation interviews, IP matrix talking heads, and short video segmented production. When using it, it is recommended to focus on controlling six items: input audio determines lip shape and duration, reference image pool determines the character for each shot, GPT-5 reverse prompts determine the action description for each segment, reference image mode determines whether to rotate, prompt mode determines whether to rotate, and strong cut mode determines the obviousness of scene cuts.
Before publishing, it is recommended to quickly fix three naming details: First, the titles of [245] / [246] are inconsistent with their actual width and height values, as they are currently 720x1280 vertical screens; second, some of images 5-10 currently appear to reuse the same placeholder image, so if you want true multi-image scene cutting, they need to be replaced with different reference images; third, you can keep LongCat15 in the output prefix, but it is recommended to add "Multi-Image Scene-Cutting Digital Human" so users can easily identify the results after downloading.
🎁Claim RH Coins first, then go experience the workflow! Avatar in top-right corner → Invitation Code → Enter [rh-v1111] to instantly claim 1000 RH Coins, and log in daily to claim another 100 coins~ For more ComfyUI workflows, tutorials, and gameplay, please follow our WeChat Official Account (AIKSK). Synchronized updates are also available on Douyin / Bilibili / Xiaohongshu / YouTube (AI-KSK) ✨


No creations yet

No creations available.