This is a two-person conversational video generation application based on LTX 2.3 + audio-driven technology. Users upload a two-person first-frame reference image and a piece of driving audio, and then fill in the two-person dialogue prompt. The system generates a continuous dialogue video centered around the two characters in the first frame, while striving to maintain stable left-right positioning, stable character identities, clear conversational relationships, and natural camera transitions. Judging from the file structure, it is not a simple single-stage talking stream, but a complete finished-video pipeline featuring three-stage sampling, two latent enlargements, audio VAE encoding/decoding, AV latent fusion, and final video composition, making it more suitable for creating high-completion two-person conversational content.
This workflow is particularly suitable for: live-action + anime dialogue, dual-character storyline voiceovers, IP character chats, light-manga two-person shots, and short dialogue videos with lip-sync. Its core selling point is not just "making images move," but: allowing the two characters in the first frame to steadily complete a continuous dialogue in fixed left and right positions. When using it, it is recommended to upload a first-frame image with clear composition and distinct positions for both people; the audio should ideally have clear vocals and a steady rhythm; and the prompts should clearly specify who is on the left, who is on the right, the scene atmosphere, and the speaking state.
🎁Claim RH Coins first, then go experience the workflow! Avatar in the upper right corner → Invitation Code → Enter [rh-v1111] to instantly get 1000RH Coins, and log in daily to get another 100 coins~ For more ComfyUI workflows, tutorials, and gameplay, welcome to follow the official WeChat account (AIKSK), with simultaneous updates on Douyin / Bilibili / Xiaohongshu / YouTube (AI-KSK) ✨
No creations yet

No creations available.