This is a video subtitle removal and local inpainting application based on LTX 2.3 Video Diffusion Model + Local Latent Reconstruction (SetLatentNoiseMask) + Audio-Video Joint Generation (AV Latent) + Two-Stage Structural Constraints (CropGuides). The system first encodes the input video, uses a Mask to control noise injection only in the specified area, and then utilizes the diffusion model to regenerate that area, achieving the effect of "removing subtitles without destroying the original video structure." Compared to traditional mosaic or overlay processing, this workflow uses the BlockifyMask + AV Latent + CropGuides structure to maintain temporal continuity, effectively avoiding flickering, edge tearing, and background collapse. It can also be combined with audio driving to keep the video rhythm consistent with the generated content.
It is particularly suitable for scenarios such as: short video subtitle removal, reposted video deduplication, ad cleanup, local video replacement, and digital human lip-sync driving generation. Its core capability is not simply "deleting subtitles," but: AI reconstruction of designated areas while keeping the original video unchanged. When using it, focus on three key elements: mask range (determines where to modify), prompt words (determines what to generate), and sampling strength (determines effect quality).
🎁Claim RH Coins first, then experience the workflow! Top right avatar → Invitation Code → Enter [rh-v1111] to instantly claim 1000 RH Coins, and log in daily to claim another 100 coins~ For more ComfyUI workflows, tutorials, and gameplay, welcome to follow our official WeChat account (AIKSK), with simultaneous updates on Douyin / Bilibili / Xiaohongshu / YouTube (AI-KSK)✨


No creations yet

No creations available.