Powered by the next-generation "Seed-Audio 1.0" model from Volcano Engine, this API supports multimodal inputs including text and multiple reference audio tracks. It features cinematic-grade multi-track mixing capabilities that automatically arrange multi-character dialogues, restore paralinguistic details (laughter, sighs, pauses), and seamlessly integrate background music and ambient sound effects in a single prompt—delivering production-ready audio without extra mixing software.