×

Microsoft launches MAI-Transcribe-2-Streaming and two MAI-Voice models – Unite.AI

Microsoft launches MAI-Transcribe-2-Streaming and two MAI-Voice models – Unite.AI

Microsoft AI on October 1, 2026 launched MAI-Transcribe-2-Streaming, its first streaming transcription template, along with two new text-to-speech templates, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, all three available through Microsoft Foundry.

MAI-Transcribe-2-Streaming and the artificial analysis benchmark

Microsoft describes MAI-Transcribe-2-Streaming as capable of providing low-latency, real-time transcriptions in 60 languages ​​with automatic, continuous language detection. The company said the model ranks first in terms of accuracy for both final and partial transcripts on artificial analysis and that it is on the Pareto frontier when evaluating accuracy versus benchmark latency, meaning that greater accuracy does not require a heavy latency trade-off. The ranking graph in the post, which cites the Artificial Analysis streaming ranking dated September 28, 2026, shows the model with a final word error rate of 2.5%, a partial prime error rate of 2.8%, and 0.13 seconds for the final transcription.

Instead of waiting for the speaker to finish before returning the text, the model produces its first guesses, known as partials, in just over 100 milliseconds of receiving the audio, then revises them as more context arrives before performing a stable transcription. Microsoft said this allows voice-enabled applications to act on speech before the speaker finishes: Voice agents can start reasoning or call tools mid-sentence, and real-time transcripts can appear as people speak. As for real-time dictation and subtitling, the company said its internal evaluations show that words appear in the transcript twice as fast as its closest competitor.

Artificial Analysis says its AA-WER Streaming Index measures transcription accuracy for models where audio is streamed in real time, piece by piece, over approximately eight hours of audio from three datasets: AA-AgentTalk at 50%, VoxPopuli at 25%, and Earnings22 at 25%. The datasets cover real-world speech with different accents, domain-specific language, and challenging acoustic conditions, and the Time to End and Time to First partial benchmark measurements both start at the end of the speech detected by the SileroVAD speech activity detector.

MAI-Transcribe-2-Streaming is available at an introductory price of $0.54 per audio hour through the end of the year. The model extends Microsoft’s MAI audio line, which already includes MAI-Transcribe-2, the previous non-streaming speech recognition model that the company said is the fastest, most accurate and cheapest in the world.

MAI-Voice-2.1 and MAI-Voice-2.1-Flash

MAI-Voice-2.1 supports 23 languages ​​and 26 locales, and Microsoft said a single voice can use them all with a native accent, maintaining the same speaker identity when switching languages. A tutoring app, in the company’s example, can switch languages ​​during class without changing teachers, and a multilingual assistant can respond in whatever language is addressed while maintaining the same voice. The template is priced at $22 for 1 million characters.

MAI-Voice-2.1-Flash supports the same languages ​​and multilingual speakers, but is designed for high-volume, latency-sensitive workloads. It can generate up to 45 seconds of audio with an end-to-end latency of 150 milliseconds, and Microsoft said it provides 55% faster model inference and is about 60% cheaper than comparable models. The price is $15 for 1 million characters.

Both voice models support voice cloning in all supported languages ​​using a few seconds of reference audio, with built-in consent barriers that Microsoft says prevent misuse. In a Turing Test of 4,000 listeners that combined the two new voice models, 50.3% of listeners rated MAI-Voice as equally or more human-like than human recordings, Microsoft said.

Microsoft defined pairing MAI-Transcribe-2-Streaming with MAI-Voice-2.1-Flash as a way to buy back time on both ends of a voice agent loop, the sequence of listening, understanding, deciding, and speaking within the window in which a human still experiences the interaction as a conversation. Developer use cases listed include customer service agents that transcribe requests as they are spoken and respond in natural language, multilingual assistants that detect spoken language and respond in any of the 23 supported MAI-Voice languages, and multimedia and interactive learning applications that use separate speakers for tutoring, role-playing, simulations, storytelling, and conversational content.

Chatter availability and demo

MAI-Voice-2.1 and MAI-Voice-2.1-Flash are available via OpenRouter. All three models are available through Microsoft Foundry, MAI Playground, Vercel, and Azure Voice Live, with LiveKit listed as coming soon.

To showcase models working together in a live agent, Microsoft created Chatter, a new demo in the MAI Playground that lets users talk to a voice assistant based on transcription and voice models.

Post Comment