The last frame has been added, and the video can now be generated in segments. Preset 10 voice tones: 0, Gentle Female Host 1, Gentle Male Broadcaster 2, Gongsun Li 3, Liu Chan 4, Conan 5, Ayumi 6, Nezha 7, Pan Jinlian 8, Ziwei 9, KFC Ad Girl
Steps:
1, Upload an image
2, Input the lines.
3, Choose the voice tone between 0 and 9. Selecting 0 allows you to upload a custom voice clone.
4, Reference audio: When choosing voice tone 0, you can upload a short speech segment to be cloned (3 seconds is sufficient, no need for long audio).
5, Speech speed: Default is 1, can be 1.1, 1.2, 0.9, 0.8
6, Speech tone and mood description: Language tone and mood can be left blank or input simply, such as "Speak in Chinese," "Speak in English," "Speak gently in Shanghai dialect," "Speak angrily in Mandarin," "Speak quickly in Mandarin," "Speak happily," "Speak sadly," "Speak angrily," "Speak gently," etc.
7, Advanced customization: In the character lines, you can input punctuation marks to control tone, such as "~", "~~", "~~~" for different tone and pause effects. You can add [laughter], for example: "I really find it so funny[laughter][laughter][laughter], I really can't hold back." A single[laughter] and three[laughter] produce different effects. This allows for fine-tuned adjustments to match the image with the voice-over.
The image uses Tencent Sonic, currently one of the best digital human technologies.