This is a digital human video generation application based on LongCat-Avatar-15 + LongCat DMD Distill LoRA + Whisper Large V3 audio lip-sync features + GPT-5 image reverse prompt generation + For-loop continuation + overlapping frame stitching + single-image soft continuation for stable talking avatars. After a user uploads a reference image of a digital human and a piece of spoken audio, the system automatically converts the audio into lip-sync features, generates LongCat-compatible talking prompt words based on the reference image, generates the first segment, continues to write subsequent segments via a For loop, and finally stitches them together into a complete talking avatar video.
The core advantage of this workflow is: generating long talking videos continuously from a single image, rather than just making a few seconds of a talking picture. In the current workflow settings, the prompt mode is 0, reference image mode is 0, and strong cut-in mode is 0, which means it will fixedly use the same character reference image and the same set of motion prompt words, using soft continuation to make subsequent segments match the previous one as much as possible. Compared to the multi-image scene-switching version, its visual changes are more restrained, but character identity, clothing, background, and camera continuity are more stable, making it suitable for formal presentations, course lecturing, news broadcasting, product introductions, virtual anchors, and brand digital human content.
It is recommended to focus on controlling six items during use: input audio determines lip-sync and video length, reference image determines character identity, user reverse-prompt needs determine motion amplitude, manual loop count determines generation length, overlapping frame count determines seam stability, and soft continuation mode determines character continuity.
The current file still retains the 10-image pool and 10-group prompt pool, but because prompt mode = 0 and reference image mode = 0, do not emphasize "multi-image scene switching" when publishing; instead, highlight single-image looping long talking avatars.
Before publishing, it is recommended to fix three places: First, the title and actual width/height of nodes [245] / [246] are inconsistent, as it is currently actually 720×1280 vertical screen; second, the output prefix still contains 10POOL—if publishing the single-image version, it is recommended to change it to LongCat15_SingleImage_Loop_Avatar; third, images 2-10 can be hidden to prevent users from mistakenly thinking they must upload multiple images.
🎁Claim RH Coins first, then go experience the workflow! Top right avatar → Invitation Code → Enter [rh-v1111] to instantly receive 1000 RH Coins, and log in daily to get another 100 RH Coins~ For more ComfyUI workflows, tutorials, and gameplay, welcome to follow the official WeChat account (AIKSK), with simultaneous updates on Douyin / Bilibili / Xiaohongshu / YouTube (AI-KSK) ✨


No creations yet

No creations available.