MultimodalCross-modal capabilities: image, video and speech
语音合成(TTS)
Technology that converts text into natural-sounding speech.
TTS powers audiobooks, voice assistants and accessibility; new-generation TTS supports voice cloning, multilingual output and emotion, billed by characters or duration.