MultimodalCross-modal capabilities: image, video and speech
语音识别(STT)
Technology that converts speech into text.
STT serves meeting transcription, subtitles and voice input; mainstream solutions support real-time streaming and speaker diarization, billed by audio duration.