MultimodalCross-modal capabilities: image, video and speech
OCR(光学字符识别)
Technology that recognizes and extracts text from images.
OCR is the base of document digitization, turning scans, screenshots and receipts into editable text. Combined with layout analysis and multimodal models, it handles complex layouts accurately.