This is a long-video consistency generation workflow based on EverAnimate + character reference image + driving video + face_video expression reference + pose_video pose control + first-segment generation + continuation loop extension. Users upload a character reference image and a driving video. The system extracts character actions, poses, facial expressions, and audio information from the driving video, and then generates a continuous video centered around the reference image.
The biggest difference between this and ordinary image-to-video is that: ordinary image-to-video usually can only generate short clips; over time, it is prone to face drifting, clothing changes, character identity changes, action breaks, and background flickering. This EverAnimate workflow splits the generation process into a connectable long-video structure through the first-segment + continuation approach.
The core value of this workflow is "reference image locks the character, driving video provides the action, and continuation loops extend the duration".
The character reference image is responsible for locking the character's identity; the driving video is responsible for providing actions, expressions, and rhythm; the first EverAnimate segment is responsible for establishing the initial stable frame; the continuation EverAnimate segments are responsible for seamlessly connecting the previous motion state; and finally, the complete MP4 video is output through the video synthesis nodes.
This workflow is more suitable for long-video character consistency generation, and is particularly suited for digital humans, talking heads, expression-driven content, half-body actions, dance movements, character presentations, storyline segments, and AI model videos.
If users want to make longer videos, this workflow is much more suitable than single-segment image-to-video because it does not forcefully stretch the video all at once, but rather generates it continuously through a continuation mechanism, making character identity, face, pose, and motion transitions much more stable.
When using it, it is recommended to focus on controlling five items: character reference image, driving video, positive prompt, number of continuation segments, and face/pose strength.
If you want the face to look more like the reference image, the face strength can be appropriately increased; if you want the action to follow the driving video more closely, the pose strength can be appropriately increased; if you want a longer video, increase the number of continuation loops, but pay attention to generation time and stability.
When publishing as a RunningHub application, the frontend is recommended to consist of:
Upload character reference image + upload driving video + fill in video description + set generation length / number of continuation segments + output video.
Model loading, VAE, LoRA, SDPose, YOLO, ViTPose, SageAttention, Torch settings, frame cutting, loop splicing, and other nodes are recommended to be fixed in the backend and not exposed to regular users.
🎁Claim RH Coins first, then go experience the workflow! Avatar in the upper right corner → Invitation Code → Enter [rh-v1111] to instantly get 1,000 RH Coins, and log in every day to get another 100 coins~ For more ComfyUI workflows, tutorials, and gameplay, please follow our WeChat Official Account (AIKSK), with simultaneous updates on Douyin / Bilibili / Xiaohongshu / YouTube (AI-KSK) ✨


No creations yet

No creations available.