Tbreak Media UAE reported that Microsoft has expanded its artificial intelligence lineup with MAI-Transcribe-2-Streaming and two new text-to-speech models, providing developers with separate components to build multilingual voice agents. The streaming transcription model supports 60 languages and returns provisional text within 100 milliseconds of audio arrival, allowing applications to display captions or trigger low-risk actions before a user finishes speaking. Conversely, the non-streaming batch model remains available at a significantly lower cost of 0.10 dollars per hour, suitable for archived recordings where immediate response is not required.

The accompanying voice models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, cover 23 languages and 26 locales, with the Flash variant offering faster generation at lower cost. Developers must compare these rates against their specific workload, as transcription is billed by audio hour while speech generation charges by character count. This modular approach separates audio-to-text conversion from reasoning and speech generation, leaving developers responsible for connecting the models and managing conversation logic.

Distribution channels include Microsoft Foundry, Azure Voice Live, and OpenRouter, with voice cloning features available for both models. While the company includes consent guardrails, developers must secure specific permissions for voice reproduction. For UAE deployments, teams must verify regional access and dialect support, as broad language counts do not guarantee quality for specific Arabic dialects. Pricing is listed in US dollars, requiring teams to confirm local availability before production design.

For more information, visit tbreak.com.