Microsoft released a streaming transcription model and two speech-generation models on 1 October 2026, adding faster ways to turn live speech into text and generate voices across languages. In its announcement of the three models, the company positioned them as components for conversational agents, live transcripts and multilingual assistants.
Key points
- Microsoft says MAI-Transcribe-2-Streaming ranks first for both partial and final transcript accuracy on Artificial Analysis.
- The transcription model detects languages automatically across 60 languages and returns its first provisional words in just over 100ms.
- MAI-Voice-2.1 supports 23 languages and 26 locales, with one voice identity able to speak across them.
- Microsoft offers a faster voice variant, MAI-Voice-2.1-Flash, at $15 per 1 million characters.
MAI-Transcribe-2-Streaming starts before speech ends
MAI-Transcribe-2-Streaming processes speech as it arrives, rather than waiting for a speaker to finish. Microsoft says it produces its first provisional transcription in just over 100ms after receiving audio, then revises those words as it hears more. A provisional word is like a pencilled-in answer: the system can put it on screen immediately, but later sounds may require a correction before the transcript becomes stable.
Dictating a sentence could bring words onto the screen before the sentence ends. Those early words could change as more speech arrives, while the developing transcript could let an application begin responding sooner.
Microsoft says the model transcribes in 60 languages and continuously detects which language is being spoken. The company reports that it ranks first for accuracy on Artificial Analysis for both partial and final transcripts. The two rankings concern different outputs: the words offered while speech is still arriving, and the settled text produced after further context. Microsoft also says its internal evaluations found that words appeared twice as fast as with its closest competitor in real-time dictation or subtitling uses.
For an application acting on speech, the provisional text creates an earlier point at which it could start work. Microsoft gives the example of a voice agent beginning to reason or call tools before a speaker has finished. For an initial offer running until the end of 2026, Microsoft charges $0.54 per hour of audio for MAI-Transcribe-2-Streaming.
MAI-Voice-2.1 keeps one voice across languages
MAI-Voice-2.1 works in the other direction, generating speech from text. Microsoft says it supports 23 languages and 26 locales, and that a single voice can speak across those languages while adopting the accent and local phrasing appropriate to each. That gives developers a way to build an assistant that switches languages without also switching its apparent speaker, according to the company.
The distinction between languages and locales matters here. A language describes what is being spoken; a locale can distinguish a regional form of it. Microsoft’s claim is therefore about more than supplying a separate voice for each language: it says one voice identity can carry through the supported range. The company prices MAI-Voice-2.1 at $22 per 1 million characters of input text.
Microsoft says both new voice models can clone a voice in their supported languages from a few seconds of reference audio. It also says they include consent safeguards intended to prevent misuse. Those provisions matter alongside the multilingual feature because a voice that can be carried from one language to another can also be based on a supplied recording.
MAI-Voice-2.1-Flash targets faster replies
MAI-Voice-2.1-Flash supports the same languages and cross-language voice identities as MAI-Voice-2.1. Microsoft describes it as the option for applications where response time and volume matter, saying it can generate 45 seconds of audio with end-to-end latency of 150ms. It reports 55% faster model inference and a cost approximately 60% below comparable models, and prices Flash at $15 per 1 million characters.
Microsoft proposes pairing Flash with the streaming transcription model in a voice agent. The transcription model could supply provisional words while someone speaks, and Flash could shorten the wait for a spoken reply. Microsoft says the time saved at those two ends of an exchange gives an agent more room to reason, use tools and check an answer during the conversation.
The release also sits alongside Microsoft’s work to help customers deploy AI. Microsoft announced in July that it would spend $2.5 billion on a unit called the Microsoft Frontier Company to help customers implement the technology, PYMNTS reported on 1 October. For the new audio models, Microsoft identifies customer service, multilingual assistance and interactive learning among the applications developers can build.
Microsoft has built a live demonstration called Chatter in its MAI Playground to show the models working together. The company says developers can access both voice models through OpenRouter, and all three models through Microsoft Foundry, MAI Playground, Vercel and Azure Voice Live.