This is an Image-to-Video application based on Bernini-R + Single Source Image + BerniniPromptEnhancer + BerniniConditioning + Two-Stage Video Sampling. Users upload an image and enter a simple video description, and the system automatically rewrites the task into video prompts optimized for Bernini-R, generating a dynamic video centered around that image.
The difference between this and standard Text-to-Video is: The image is the visual starting point, while the text serves as action and plot instructions.
The source image is responsible for locking in the characters, scene, composition, visual style, and initial state; the prompts tell the model what happens next, such as facial expression changes, camera movement, background events, action reactions, scene atmosphere, and dynamic rhythm. This makes it much easier to keep the main subject clear compared to pure Text-to-Video, and it is better suited for short video shots, animated covers, and character dynamic performances.
The current default case is a scene well-suited for showcasing the model's dynamic capabilities: A woman by the racetrack holding a camera, watching an F1 car and a black truck in an intense race as the vehicles speed past, kicking up smoke, while her expression shifts from normal to intense shock.
This case demonstrates several key capabilities of Bernini-R Image-to-Video: facial expression changes, racetrack environment dynamics, high-speed vehicle motion, smoke effects, foreground-background relationships, and cinematic impact.
The core advantage of this workflow is "One image + One action description → Generate a complete dynamic shot".
Users do not need to prepare raw videos or multiple reference images. With just a single image, they can expand a static picture into a video complete with actions, emotions, and scene events. It is ideal for AI short video assets, character reaction videos, product dynamic displays, storyboard scenes, cinematic shots, cyber/fantasy/realistic scene motion effects, Xiaohongshu cover videos, Bilibili animated covers, and RunningHub beginner-friendly Image-to-Video applications.
When using, it is recommended to focus on controlling three items: The input image determines the visual foundation, the Image-to-Video prompt determines the action and plot, and the video frame count determines the final video duration.
If you want more stable results, the prompt should explicitly state: maintain character identity consistency, keep the original composition, avoid overly drastic actions, keep the background stable, make camera movements slow or explicit, do not add unrelated subjects, and avoid text watermarks. If you want stronger dynamic effects, you can clearly specify event changes, such as "vehicles speeding past from the right", "hair and clothes blowing in the wind", or "the character freezes first then widens their eyes".
When publishing as a RunningHub application, the frontend is recommended to have a minimalist structure: Upload image + Fill in video description + Set video length + Output video. Sampling steps and staged sampling cut points are recommended to be fixed in the backend to prevent user misadjustments leading to image drift, character deformation, or video flickering.
🎁Claim RH Coins first, then go experience the workflow! Top right avatar → Invitation code → Enter [rh-v1111] to instantly claim 1000 RH Coins, and log in daily to claim another 100 coins~ For more ComfyUI workflows, tutorials, and gameplay, follow our official WeChat account (AIKSK), with simultaneous updates on Douyin / Bilibili / Xiaohongshu / YouTube (AI-KSK) ✨
No creations yet

No creations available.