MultimodalCross-modal capabilities: image, video and speech

视觉语言模型(VLM)

Models combining a vision encoder with a language model for image understanding.

VLMs convert images into visual tokens for the language model, handling image captioning, chart reading and screenshot analysis. GPT-image, Gemini and Qwen-VL belong here.

Related terms