MultimodalCross-modal capabilities: image, video and speech
图像理解
A model's ability to recognize and explain image content.
Image understanding covers object recognition, scene description, chart parsing and document OCR—the core input skill of multimodal models, widely used in auto-tagging, content moderation and accessibility.