MultimodalCross-modal capabilities: image, video and speech

文生视频

AI that generates video clips from text descriptions.

Text-to-video generates coherent frames via diffusion or autoregressive frameworks; mainstream models (Sora, Seedance, Kling, Hailuo) bill per second and resolution—among the priciest generation workloads.

Related terms