News

Microsoft’s MAI-Transcribe-2 Beats OpenAI and Google at Just $0.10 Per Hour

4 min read Bhavesh

Microsoft has officially introduced MAI-Transcribe-2, a new in-house speech recognition model the company says is built to set fresh industry standards across speed, accuracy, and cost-efficiency. In a short announcement, Microsoft positioned the model as a direct challenger to offerings from OpenAI and Google — and pointed to a striking price tag as its headline differentiator: just $0.10 per hour of transcribed audio.

Here’s what we know about the new model, how it compares to rivals, and what it could mean for developers and businesses that rely on turning speech into text.

What MAI-Transcribe-2 is

MAI-Transcribe-2 is Microsoft’s latest speech recognition model developed in-house, according to the company’s announcement. Speech recognition — sometimes called automatic speech recognition, or ASR — is the technology that converts spoken language into written text, powering everything from live captioning and meeting transcripts to voice search and accessibility features.

Advertisement

The “MAI” prefix ties the model to Microsoft’s broader AI brand, which has been used for a range of in-house machine learning products. The “2” in the name suggests this is a second generation, implying the company is iterating on an earlier transcription model rather than launching the category from scratch. That framing places it in a lineage that stretches back through years of Microsoft speech work, from early dictation tools in Windows to later voice-driven features.

Microsoft says the model targets three benchmarks at once: it aims to be faster at processing audio, more accurate in what it transcribes, and cheaper to run than competing options. Those three goals often pull in opposite directions in AI model design, so hitting all three is what the company is betting will set it apart.

A split-screen composition showing an audio waveform on the left and transcribed text lines on the right, rendered in a
Speech recognition models like MAI-Transcribe-2 convert spoken audio into written text in real time.

How it compares to OpenAI and Google

The announcement explicitly names OpenAI and Google as the benchmarks MAI-Transcribe-2 is meant to beat. Both companies run large speech recognition systems that power widely used products — OpenAI’s Whisper model, for instance, became one of the most popular open-source transcription tools available, while Google’s speech tech underpins features across Android, Pixel devices, and its search and Assistant products.

Microsoft’s claim is that its new model outperforms those rivals on accuracy and speed while costing far less. The company did not release detailed benchmark tables in the summary of the announcement, so the exact margins of the lead are not independently verified — treat the head-to-head numbers as Microsoft’s own figures until third-party testing confirms them.

What is notable is the framing: rather than competing purely on raw accuracy, Microsoft is pitching a combination of quality and price. In an era where AI transcription costs can add up quickly for businesses processing large volumes of audio, a model that claims to match top-tier accuracy at a fraction of the cost is a compelling pitch to developers weighing their options.

The $0.10-per-hour pricing angle

Perhaps the most attention-grabbing detail is the price: $0.10 per hour of audio transcribed. To put that in context, that works out to roughly a tenth of a cent per minute, which is well below what many commercial transcription services and API-based speech models charge.

For a business or developer processing, say, ten hours of audio a day, that translates to a modest daily cost — the kind of metric that matters when transcription is baked into a product serving thousands of users. Lower per-hour costs can make previously expensive features, like automatic meeting summaries or searchable voice recordings, economically viable at scale.

Again, the $0.10 figure comes directly from Microsoft’s announcement. It is worth confirming whether that rate applies uniformly or varies by language, audio quality, or volume tier before you build anything on top of it.

What this means for you

If you’re a developer or product manager evaluating speech-to-text options, MAI-Transcribe-2 gives you another candidate to test — one pitched specifically on the combination of quality, speed, and price. The practical question is whether Microsoft’s own benchmarks hold up against independent testing, especially across different accents, background noise, and domain-specific vocabulary.

For everyday Windows users, the broader signal is that competition among big AI labs is pushing transcription quality up and prices down. That trend tends to filter down into consumer features over time — better live captioning, smarter voice commands, and richer accessibility tools in products you already use.

Microsoft hasn’t yet confirmed a broad rollout date or exactly which products will ship with MAI-Transcribe-2 first. Until then, treat it as an early-stage model announcement rather than a finished feature sitting in your next update.

How to find out more

Microsoft confirmed the unveiling of MAI-Transcribe-2 through its own announcement channels. Developers interested in testing the model should watch for official documentation and API access details from Microsoft, which typically follow the initial announcement with integration guides and pricing pages.

Until those details land, the safest move is to hold off on committing to the model for production use and instead track Microsoft’s follow-up posts for benchmark data, language support, and access instructions.

Source: Neowin

Over to you: Would a $0.10-per-hour transcription model change which speech-to-text service you choose for your projects?

Advertisement
Share:
Bhavesh
Written by
Bhavesh

Tech journalist covering Windows, Microsoft, and PC hardware. Bhavesh has followed the Windows ecosystem since Windows 7 and writes with a focus on practical user impact and technical accuracy.

Advertisement