Microsoft once relied nearly entirely on OpenAI's frontier models to power its cutting-edge tools. But now it's increasingly launching its own models.
On Thursday, Microsoft unveiled its first-ever streaming transcription model, MAI-Transcribe-2-Streaming, offering low-latency, real-time transcripts in 60 languages. At launch, it topped the Artificial Analysis Word-Error-Rate Leaderboard, which measures the percentage of words transcribed incorrectly in the final transcription, beating out models from Grok, Meta, and OpenAI.
Microsoft claims that the model can deliver text as soon as 100 milliseconds after receiving audio. In AI applications, it allows models to begin reasoning sooner, enabling them to begin tasks before the user has finished their thought.
The model also helps with low-latency, real time subtitling. Microsoft claims that internal evaluations found it to be twice as fast as its closest competitors for use cases such as real-time dictation or subtitling.
Additionally, Microsoft launched two new voice models:
- MAI-Voice-2.1: Strongest multilingual text-to-speech model, with expanded support to 23 languages and 26 locales.
- MAI-Voice-2.1-Flash: Has the same language support as the model above, but has also been optimized for high-volume, latency-sensitive workloads.
As voice models start to power seamless interactions with voice agents across AI software and hardware, Microsoft isn't the only company with its eye on the space. Suno, best known for its music-generation models, has now expanded beyond its core offering with the release of Speech, which it calls the first audio model to generate voice and music in a single, cohesive track, the company announced on Thursday.
The way Suno Speech works is that users can create spoken audio set to background music by typing text, describing the voice and musical style, and then the model outputs a reading in a voice that fits the user's description with original music to go along with it. It is available in beta on the platform, with the company caveat that, since it is still in beta, it doesn't always perform accurately, for instance, noting that British accents drift into Australian, and dramatic pauses may sound overdone.
Our Deeper View
Microsoft's latest model releases are all domain-specific: MAI-Thinking-1, Microsoft AI's first reasoning model; MAI-Code-1.1-Flash, a model focused on coding assistance; and MAI-Image-2.6, its most advanced image model. Microsoft has several reasons to stick with this strategy. By focusing each model on a single domain, it can concentrate its resources on making that model high-quality and high-performing, as these smaller models' strong benchmark results show. Building general models that cover a broad range of use cases is very resource-intensive. The approach also suits Microsoft's business and enterprise customers, where the company has its best opportunity to win since enterprise users rely on AI for specific tasks in their everyday workflows. For example, for an engineer, a cutting-edge coding model matters more than a general model that can tell them weather patterns and do math.




