MultimodalCross-modal capabilities: image, video and speech

语音合成(TTS)

Technology that converts text into natural-sounding speech.

TTS powers audiobooks, voice assistants and accessibility; new-generation TTS supports voice cloning, multilingual output and emotion, billed by characters or duration.

Related terms