MultimodalCross-modal capabilities: image, video and speech
视觉语言模型(VLM)
Models combining a vision encoder with a language model for image understanding.
VLMs convert images into visual tokens for the language model, handling image captioning, chart reading and screenshot analysis. GPT-image, Gemini and Qwen-VL belong here.