Microsoft AI releases MAI-Transcribe-2-Streaming, tops Artificial Analysis real-time speech-to-text ranking
Microsoft AI launched MAI-Transcribe-2-Streaming, a real-time speech-to-text model ranked first by Artificial Analysis.
MAI-Transcribe-2-Streaming is the real-time sibling of the batch MAI-Transcribe-2, which was released in September. It transcribes 60 languages with automatic, continuous language detection. Audio streams in continuously and text streams back while the speaker is still talking. The model emits its first hypotheses, called partials, just over 100 milliseconds after receiving audio, revises them as context arrives and then commits a stable final transcript. According to Microsoft's internal tests, words appear twice as fast as with its closest competitor. For an agent, that means it can start reasoning or calling tools mid-sentence.
Artificial Analysis measured the model with its AA-WER Streaming index, using about eight hours of audio. The mix is AA-AgentTalk at 50%, VoxPopuli at 25% and Earnings22 at 25%. Latency is timed from the end of speech as detected by SileroVAD. MAI-Transcribe-2-Streaming recorded 2.5% word error rate for final transcripts at 0.13 seconds after the end of speech, ranking first of 38 models. Its first partial transcript also scored 2.5% WER at 0.12 seconds, also first. Grok Voice Transcribe 2.0 was second at 2.7% WER and 0.49 seconds, while Muse Voice Transcribe reached 3.1% WER at 0.16 seconds. Cartesia Ink-2 was faster on external endpoints, returning finals in 0.07 seconds, but at 4.0% WER. Artificial Analysis said the first partial is as accurate as the final transcript, a detail that matters for agents acting before a speaker finishes. Microsoft also placed the model on the accuracy versus latency Pareto frontier.
Pricing for MAI-Transcribe-2-Streaming is $0.54 per hour of audio, an introductory price through the end of 2026. Artificial Analysis normalized that to $9.00 per 1,000 minutes. The batch MAI-Transcribe-2 costs $0.10 per hour. In streaming, Microsoft charges more than xAI and Meta and roughly matches Google's estimated rate. The model is in public preview with no service-level agreement and no open weights.
Microsoft documents two integration paths. The Realtime API suits applications already using an OpenAI Realtime-compatible WebSocket, while the Azure Speech SDK handles connection management, retries and audio streaming. Both return intermediate and final results. The model is also available in the MAI Playground, through Vercel and Azure Voice Live. LiveKit support is listed as coming soon. Microsoft pairs it with MAI-Voice-2.1-Flash for full voice loops; Flash generates 45 seconds of audio at 150 milliseconds end-to-end latency for $15 per 1 million characters. MAI-Voice-2.1 covers 23 languages and 26 locales at $22 per 1 million characters.
Artificial Analysis' comparison lists MAI-Transcribe-2-Streaming against Grok Voice Transcribe 2.0, Muse Voice Transcribe and Gemini 3.5 Transcribe Live. The four models were released between August 26 and October 1, 2026. Their streaming prices are $0.54 per hour for Microsoft's introductory rate, $0.20 for Grok and $0.18 for Muse, with Google's token-billed rate estimated at about $0.54 per hour. None of the four offers open weights. Speaker diarization in the stream is not stated for MAI-Transcribe-2-Streaming, while Muse supports it and Gemini does not support it in Live mode. Microsoft's model supports 60 languages with continuous automatic detection; Muse lists more than 70 trained languages with 25 verified, and Gemini lists more than 85 with auto-detect. Grok lists dozens with automatic detection and mid-recording switching.