MultimodalCross-modal capabilities: image, video and speech

声音克隆

Replicating a specific voice's timbre and style from a small sample.

Voice cloning uses speaker embeddings and conditional generation to synthesize similar voices from brief recordings, powering audiobooks and personalized assistants—but also deepfakes, requiring consent and detection safeguards.

Related terms