MultimodalCross-modal capabilities: image, video and speech
文生图
AI that generates images from text descriptions.
Text-to-image is dominated by diffusion models: starting from noise, iteratively denoising toward the described image. Stable Diffusion, FLUX, Midjourney and Seedream bill per image by resolution—priced in this site's image section.