This is a reference video conditional editing application based on Bernini-R + original video + reference video + BerniniConditioning + two-stage video editing sampling. Users upload an original video, upload a reference video, and enter an editing requirement. The system embeds the reference video content into designated scenes of the original video based on the prompt, generating a newly edited video.
The biggest difference between this and ordinary video generation is: It does not generate videos from scratch, nor does it just use a single reference image for face-swapping or background replacement. Instead, it integrates the "reference video" as propagatable video conditional content into the main video editing.
The original video provides the main scene, camera structure, and final image base; the reference video provides the dynamic content to be implanted, displayed, or propagated; the editing prompt determines how the reference video should appear in the main video, such as on a billboard, TV screen, phone screen, stage big screen, car screen, LED screen, store window screen, or background projection.
The current default case is: Putting the reference video as an F1 racing championship video to loop on a street billboard. This direction is very suitable for demonstrating Bernini-R's "video-to-video conditional editing" capability, because it is not single-frame texture mapping, but requires the dynamic content of the reference video to maintain temporal consistency in the new scene, while blending with the spatial, perspective, lighting, and camera relationship of the original video.
The core advantage of this workflow is "main video scene + reference video content + conditional editing fusion".
Traditional compositing methods require manual screen keying, tracking, perspective distortion, color grading, occlusion handling, and lighting adjustment; whereas this type of reference video conditional editing workflow can leave these steps to model generation, allowing the reference video to naturally enter the original video frame. It is particularly suitable for ad insertion, screen content replacement, video footage embedding, brand exposure, sports event footage implantation, TV screen replacement, street big screen replacement, virtual billboards, derivative storytelling videos, and commercial short video revisions.
When using, it is recommended to focus on controlling three items: The original video determines the main scene, the reference video determines the implanted dynamic content, and the editing prompt determines the location and method of the reference video's appearance.
For better stability, the prompt should clearly specify: which carrier the reference video appears on, whether it loops, whether the original video scene remains unchanged, whether characters or cameras remain unchanged, whether the reference video needs to fit the screen perspective, and whether natural lighting fusion is required.
When publishing as a RunningHub application, the frontend is recommended to be designed as: Upload original video + Upload reference video + Fill in editing requirements + Output video. Video frame count can be opened as a basic parameter; sampling steps and stage sampling cut points are recommended to be fixed in the backend to prevent users from misadjusting them, which could cause reference video fusion failure, scene destruction, or screen flickering.
🎁Claim RH Coins first, then experience the workflow! Avatar in the upper right corner → Invitation Code → Enter [rh-v1111] to instantly claim 1000 RH Coins. Log in daily to claim another 100 RH Coins~ For more ComfyUI workflows, tutorials, and gameplay, follow our official WeChat account (AIKSK). Synchronized updates also available on TikTok / Bilibili / Xiaohongshu / YouTube (AI-KSK) ✨
No creations yet

No creations available.