This is a single-person digital human video application based on LTX2.3 Dev-Dare TIES Distilled 1.1 + Single-Image Character Reference + Audio-Driven + MelBandRoFormer Vocal Separation + PromptRelay Global/Segmented Prompts + VBVR Temporal Stability + OmniNFT Enhancement + sULph LoRA + LTX2_NAG + Third-Order Sampling + 10S Face Similarity Maintenance System. After users upload a character image and a voiceover or singing audio clip, the system automatically trims the audio, extracts vocals, reads the audio duration, calculates video frames, and makes the character naturally speak or sing to the rhythm of the voice while trying to maintain consistency in the character's face, hairstyle, clothing, background, and identity.
The biggest upgrade in this version is 10S Similarity Maintenance. Rather than simply relying on image-to-video strength to forcefully adhere to the reference image, it introduces face detection, similarity guidance, similarity anchors, latent anchor perception, STG advanced guidance, and Sigmas Easing. The overall strategy is: first-order strong similarity establishes the character identity, second-order medium similarity maintains high-definition upscaling without face-swapping, and third-order weakened anchors preserve the naturalness of lip movements and expressions. This is better suited for generating around 10-second talking-head, singing, and character-consistency videos compared to mere I2V strength.
Suitable scenarios include: AI talking heads, character singing, virtual anchors, digital human explanations, IP character voicing, anime character lip-syncing, advertising voiceovers, short video e-commerce, MV character clips, and RunningHub single-person digital human applications. When using, it is recommended to focus on controlling five key aspects: the character's first-frame image determines the character identity, the driving audio determines lip shape and rhythm, global prompts (used to write overall visual rules) determine visual stability rules, segmented prompts (writing action changes for each time period) determine expressions and actions, and similarity maintenance strength determines how closely it resembles the original image.
Before publishing, it is recommended to revise four points: First, the [520] / [521] node titles still read "Single-Person Dialogue", which should be uniformly changed to your currently fixed "Global Prompts (used to write overall visual rules)" and "Segmented Prompts (writing action changes for each time period)"; second, there are currently multiple SaveVideo preview outputs, and it is recommended to keep only the third-order final output [510] on the frontend; third, the OmniNFT node title reads 0.65 but the actual strength is about 1.0, so it is recommended to unify the title and value after actual testing; fourth, IC Subtitle Removal / Watermark Removal LoRA is currently bypassed, so do not promote "IC Edit Subtitle and Watermark Removal" as a core feature unless the cleanup branch is re-enabled.
🎁Claim RH Coins first, then go experience the workflow! Avatar in the upper right corner → Invitation Code → Enter [rh-v1111] to instantly receive 1000 RH Coins, and log in daily to get another 100 RH Coins~ For more ComfyUI workflows, tutorials, and gameplay, please follow our WeChat Official Account (AIKSK), with simultaneous updates on Douyin / Bilibili / Xiaohongshu / YouTube (AI-KSK) ✨
No creations yet

No creations available.