Microsoft adds streaming transcription and voice models to MAI family
Microsoft introduced MAI-Transcribe-2-Streaming and two MAI-Voice-2.1 text-to-speech models for developers building near-instant voice agents, with streaming transcription priced at $0.54 per audio hour.
MAI-Transcribe-2-Streaming accepts human speech through a WebSocket and turns it into a transcript that is continuously updated as a person keeps talking, Microsoft said. When the speaker finishes, the model confirms that the transcript is final. The company said this allows applications to display live captions or begin processing a user's request before the user has finished speaking. The model is listed on Microsoft's Vercel AI Gateway at $0.54 per audio hour. It supports more than 60 languages and can automatically detect which language is being spoken. Microsoft said it returns first transcript hypotheses within 320 milliseconds on average, though it cannot guarantee that performance in every scenario because response speed also depends on network connection and the AI system generating the reply.
The streaming model costs more than Microsoft's earlier transcription release. Last month, the company debuted MAI-Transcribe-2 at $0.10 per audio hour, more than five times cheaper than the streaming variant. Microsoft attributed the difference to the greater processing required by the streaming model, which returns provisional results while audio is still arriving. The regular MAI-Transcribe-2 waits for the speaker to finish before processing begins.
The two text-to-speech models address the other side of conversational AI. MAI-Voice-2.1 is intended for more expressive, higher-fidelity output, while MAI-Voice-2.1-Flash trades some of those qualities for faster response and lower cost. According to Vercel AI Gateway listings, MAI-Voice-2.1 costs $22 per million characters and MAI-Voice-2.1-Flash costs $15 per million characters. Both support 23 languages, Microsoft said.
The release also reflects Microsoft's effort to rely more on its own MAI models. Microsoft is a major investor in OpenAI Group PBC and Anthropic PBC, but in July it was reported that Microsoft AI CEO Mustafa Suleyman was increasingly concerned about the cost of using those companies' frontier models. He instructed Microsoft AI researchers to increase work on the MAI family, with the eventual goal of using those models to power Copilot agents in platforms such as Excel and Outlook. "We pay a lot of money to Anthropic, so our goal is to reduce and ultimately eliminate that cost," Suleyman told Bloomberg in an interview.
Microsoft said the new models give developers the components needed to build voice AI agents that can converse in a natural, humanlike way. Such agents need to recognize and understand what a person says, decide what to do based on that speech, and generate an audible response. MAI-Transcribe-2-Streaming handles the first task and the MAI-Voice models handle the third. In the middle is Microsoft's reasoning model, Mai-Thinking-1, a standard large language model that reads transcripts to decide what the agent should do. By using three separate models, Microsoft said developers have more control over the quality, latency and costs of their voice agents.
Editor's Summary
Microsoft expanded its MAI model family with MAI-Transcribe-2-Streaming and two MAI-Voice-2.1 text-to-speech models, giving developers tools for faster, more natural voice agents. The streaming transcription model costs $0.54 per audio hour, supports more than 60 languages and returns initial hypotheses in about 320 milliseconds on average. The release also signals Microsoft's push to reduce reliance on OpenAI and Anthropic models by expanding its in-house MAI lineup for Copilot and other products.