MultimodalCross-modal capabilities: image, video and speech

文生图

AI that generates images from text descriptions.

Text-to-image is dominated by diffusion models: starting from noise, iteratively denoising toward the described image. Stable Diffusion, FLUX, Midjourney and Seedream bill per image by resolution—priced in this site's image section.

Related terms