MultimodalCross-modal capabilities: image, video and speech
文生视频
AI that generates video clips from text descriptions.
Text-to-video generates coherent frames via diffusion or autoregressive frameworks; mainstream models (Sora, Seedance, Kling, Hailuo) bill per second and resolution—among the priciest generation workloads.