MultimodalCross-modal capabilities: image, video and speech

图像理解

A model's ability to recognize and explain image content.

Image understanding covers object recognition, scene description, chart parsing and document OCR—the core input skill of multimodal models, widely used in auto-tagging, content moderation and accessibility.

Related terms