Microsoft ships a live transcription model at $0.54 an hour
MAI-Transcribe-2-Streaming returns partial text in about 100 milliseconds, and two new voice models join it for real-time agents.

Key takeaways
- Microsoft's first streaming transcription model, MAI-Transcribe-2-Streaming, supports 60 languages and returns provisional text in just over 100 milliseconds.
- It launches at an introductory $0.54 per hour of audio through the end of 2026, against $0.10 an hour for Microsoft's non-streaming model shipped in September.
- MAI-Voice-2.1 covers 23 languages and 26 locales at $22 per million characters; the Flash variant costs $15 and can generate 45 seconds of audio at 150 milliseconds of latency.
- Both voice models support cloning from a few seconds of reference audio and include consent guardrails.
Microsoft has shipped the first streaming model in its MAI family, alongside two new text-to-speech systems, in a release aimed squarely at developers building voice agents that can answer while a person is still talking. MAI-Transcribe-2-Streaming accepts audio over a WebSocket, returns provisional text as the words arrive, and revises it as more context comes in before committing a final transcript, the company said in its launch post.
The headline figure is latency. Microsoft says the model produces its first partial hypotheses in just over 100 milliseconds of receiving audio, and that it ranks first on Artificial Analysis for accuracy on both partial and final transcripts. It supports 60 languages with automatic, continuous language detection, so a conversation that switches language mid-sentence does not have to be restarted. SiliconANGLE puts the average time to a first hypothesis at 320 milliseconds.
The price is where the trade-off shows. Streaming costs $0.54 per hour of audio as an introductory rate through the end of 2026. Microsoft's non-streaming MAI-Transcribe-2, released in September, runs $0.10 per hour through the same date, because it waits for the speaker to finish rather than processing as they go.
The two voice models run the other direction. MAI-Voice-2.1 is Microsoft's most expressive multilingual text-to-speech model, covering 23 languages and 26 locales, and it keeps one speaker identity across all of them rather than dragging a single accent from language to language. It is priced at $22 per million characters. MAI-Voice-2.1-Flash keeps the same language coverage and adds speed: 150 milliseconds of end-to-end latency, 45 seconds of audio per generation, and a price of $15 per million characters, which Microsoft says is about 55 percent faster and 60 percent cheaper than comparable models. Both support cloning a voice from a few seconds of reference audio, with consent guardrails the company says are built in.
There is a caveat worth carrying. Windows Forum notes that Microsoft Learn lists the transcription and voice features as public previews without a service-level agreement, and does not recommend them for production workloads. Availability is not the same as production readiness, and the introductory price runs only to the end of the year.
For video makers, the practical piece is the voice stack rather than the transcription. A multilingual voice that holds its identity while switching languages is the thing that makes a dubbed cut sound like one narrator rather than several, and $15 per million characters is cheap enough to test at volume. The models are available through Microsoft Foundry, the MAI Playground, Vercel and Azure Voice Live, with the two voice models also on OpenRouter.
The three models are live; Microsoft has not said what the streaming transcription price becomes after December 31.
Sources
- microsoft.ai - the launch, the specifications and the prices
- siliconangle.com - the WebSocket mechanism and the earlier batch model's price
- heise.de - the Artificial Analysis placement and the language switching
- windowsforum.com - the public-preview status and the absence of an SLA