This is an audio asset generation application based on Stable Audio 3 Medium Base + T5Gemma Stable Audio text encoder + categorized intelligent expansion + Music / Instrument / SFX / One-shot four sound templates + KSampler audio sampling + VAE audio decoding. Users only need to input a simple sound prompt, select the audio type and duration, and the system can generate corresponding music, instrument materials, sound effects, ambient sound, or one-shot sound assets.
The core advantage of this workflow is: Instead of just throwing a short prompt directly to the model, it first rewrites the prompt according to the sound asset type. If the user selects Music, it expands into a complete music prompt containing style, instruments, rhythm, mood, BPM, and duration; if Instrument is selected, it generates prompts suitable for loops / stems / instrument segments; if SFX is selected, it emphasizes the sound source, material, space, movement, and temporal evolution; if One-shot is selected, it generates short, isolated single sounds usable for music production or game interaction. It is suitable for short drama soundtracks, game sound effects, video transition sounds, button prompts, ambient atmosphere sounds, instrument loops, BGM drafts, film and television sound effects, advertising sound materials, and AI video soundtrack assets. When using, it is recommended to focus on controlling four items: user sound prompt determines the generation direction, audio type determines the expansion logic, audio duration determines the material length, and random seed determines reproducibility.
🎁Claim RH Coins first, then experience the workflow! Avatar in the upper right corner → Invitation Code → Enter [rh-v1111] to instantly get 1000 RH Coins, and log in daily to get another 100 RH Coins~ For more ComfyUI workflows, tutorials, and gameplay, welcome to follow the official WeChat account (AIKSK), and TikTok / Bilibili / Xiaohongshu / YouTube (AI-KSK) with synchronized updates ✨


No creations yet

No creations available.