News
Microsoft’s Cheaper AI Now Handles 28 Million Patient Visits
Microsoft’s new MAI models promise up to 89% lower GPU costs, but its biggest deployment now transcribes patient visits with no outside safety audit.
Microsoft AI pushed two new in-house models into public preview on Wednesday, MAI-Image-2.5-Pro and MAI-Voice-2-Flash, alongside internal data claiming GPU cost cuts of up to 89% versus OpenAI’s models, VentureBeat reported. Both models now run in production across Bing, PowerPoint, OneDrive, Dynamics 365, Excel, GitHub Copilot and Azure.
But the most striking figure in that data dump is 28 million, the patient encounters a Microsoft speech model now handles every quarter, graded so far by nobody but Microsoft itself.
Two Models, Built for Opposite Ends of the Cost Curve
The two releases sit at opposite ends of what Microsoft calls its quality-speed-cost curve, and the split is deliberate.
MAI-Image-2.5-Pro targets the premium tier: hero imagery, detailed editing and the kind of in-image text rendering that has tripped up nearly every image generator on the market. Microsoft priced it at $5 per million text input tokens, $8 per million image input tokens and $106 per million image output tokens. The base MAI-Image-2.5 model already ranked No. 2 for image editing on Arena, the community leaderboard that doubles as a scoreboard for generative media.
Rob Reilly, global chief creative officer at advertising giant WPP, called the Pro model “a strong leap forward for GenMedia tools” in a statement Microsoft included in its own announcement, adding that “Microsoft has firmly established itself among the leaders in generative AI.”
MAI-Voice-2-Flash runs the opposite play. First shown at Microsoft’s Build conference, it processes speech twice as fast as MAI-Voice-2 and costs 32% less, priced at $15 per million characters. It’s built for the unglamorous, high-volume end of voice AI: call centers, voice agents and real-time speech tools, where latency and cost per call matter more than expressive range.
| Model | Built For | Pricing | Efficiency Claim |
|---|---|---|---|
| MAI-Image-2.5-Pro | Hero imagery, detailed editing, in-image text | $5 per million text input tokens; $8 per million image input tokens; $106 per million image output tokens | Base model ranked No. 2 for image editing on Arena |
| MAI-Voice-2-Flash | Call centers, voice agents, real-time speech | $15 per million characters | Twice the speed of MAI-Voice-2 at 32% lower cost |
Neither model is the headline, though. The headline sits in the deployment data Microsoft chose to publish alongside them.
A Speech Model Lands Inside 28 Million Patient Visits
Buried in that deployment list is the one that carries the most consequence: Microsoft’s Dragon Copilot, the ambient documentation tool built into clinical workflows across the country, just moved its multilingual transcription onto MAI-Transcribe-1.5.
This isn’t Microsoft’s first attempt to prove Dragon Copilot works. When the product launched, the company backed it with a survey of 879 clinicians across 340 organizations and a separate survey of 413 patients, both commissioned by Microsoft itself, detailed in a March 2025 healthcare blog post. The error-rate figures Microsoft published this week follow the same pattern: an internal number, measured on an internal test set, published without an outside audit.
What we know:
- Dragon Copilot serves 170,000 medical providers and processed 28 million patient encounters last quarter.
- MAI-Transcribe-1.5 now runs its multilingual workflow across 58 languages.
- Microsoft’s internal evaluations report a 50% relative reduction in transcription and language-identification error rates across most of those languages.
What’s unconfirmed:
- No independently published, peer-reviewed clinical accuracy study of MAI-Transcribe-1.5 exists yet.
- Whether the improvement holds evenly across all 58 languages, or clusters in the handful Microsoft tests most heavily, hasn’t been broken out.
- How the model’s real-world error rate compares with rival clinical transcription tools audited by outside researchers is unknown, since Microsoft picks its own comparison set.
None of that makes the 50% figure false. It just makes it a claim from one side of the table, in a workflow where a mistranscribed dose or symptom can end up in a permanent chart.
Does Anyone Outside Microsoft Get to Check the Transcript?
Not yet. No independent lab has published a peer-reviewed accuracy study of MAI-Transcribe-1.5. Microsoft’s error-reduction figures come from its own internal evaluation, measured against its own test set, graded by its own team, the same setup that has validated Dragon Copilot’s marketing claims since the product first launched.
The closest outside scrutiny available covers a different company’s product, and it isn’t reassuring. In 2024, researchers from Cornell University and the University of Washington found that OpenAI’s Whisper transcription model fabricated content, sometimes violent or racially charged, in about 1.4% of the transcriptions they examined. Separate peer-reviewed research has traced some of this behavior to how ASR systems hallucinate text from non-speech audio, inventing sentences from silence or background noise.
One Whisper-based tool, built by a company called Nabla, was already in use by more than 30,000 clinicians across dozens of health systems, transcribing millions of medical visits, according to reporting on the study. The tool deleted the original audio recording, leaving clinicians no way to check the transcript against what a patient actually said. OpenAI itself had warned against using Whisper in “high-risk domains.” Hospitals used it anyway.
That particular scandal belongs to OpenAI, not Microsoft. But weeks before this announcement, ECRI Institute, an independent nonprofit that has published its hazard ranking since 2008, singled out Microsoft’s own brand in a different report. In January, ECRI ranked misuse of AI chatbots as the single biggest health technology hazard of 2026.
Rather than truly understanding context or meaning, AI systems generate responses by predicting sequences of words based on patterns learned from their training data. They are programmed to sound confident and to always provide an answer to satisfy the user, even when the answer isn’t reliable.
ECRI executives wrote in the report, which ranked chatbot misuse the top health hazard of 2026 and named Copilot among the large language models clinicians and patients increasingly lean on, alongside ChatGPT, Claude, Gemini and Grok. The report wasn’t a review of Dragon Copilot’s transcription pipeline specifically. But it shows the same Microsoft brand, and the same underlying MAI model family, drawing scrutiny from patient-safety researchers in the very month Microsoft chose to expand it.
The Rest of the Ledger Holds Up Better
Outside the hospital, Microsoft’s efficiency numbers are easier to take at face value, mostly because a bad output there means a re-generated slide, not a clinical note.
- Bing Image Creator now runs entirely on MAI-Image-2.5, the first time Microsoft’s consumer image tool has gone fully in-house end to end.
- PowerPoint sees GPU costs cut by up to 84% compared with GPT-Image-2, OpenAI’s image model.
- OneDrive uses MAI-Image-2.5 as the default for key editing scenarios, with a 26% increase in save rates, roughly 25% lower P95 latency and 2.5 times the efficiency under medium-utilization workloads.
- Dynamics 365 Contact Center, used by customers including T-Mobile and EasyJet, runs MAI-Voice-2-Flash for GPU cost reductions Microsoft puts as high as 89%.
- GitHub Copilot runs MAI-Code-1-Flash, which Microsoft says posts roughly a 10% higher code accept rate than GPT-5.4 Mini and Claude Haiku 4.5 in VS Code, using 10% fewer median tokens.
Developer loyalty tracked the accept-rate numbers. Microsoft says users were 6% more likely to return across multiple days with MAI-Code-1-Flash than with GPT-5.4 Mini, and 11% more likely than with Claude Haiku 4.5.
Then Microsoft pushed the same checkpoint further. It took MAI-Code-1-Flash and trained it inside an Excel reinforcement-learning environment, teaching a coding model the specific workflows of spreadsheet work. Production feedback puts the result on par with GPT-5.6 for the most common Excel tasks, while running on Nvidia’s older H100 and even A100 chips instead of the newest accelerators.
That hardware detail carries its own weight. Every major AI company is fighting for allocation of Nvidia’s newest chips, and a model that matches frontier quality on two-generation-old silicon changes the deployment math. It also frees Microsoft’s newly operational GB200 cluster for training instead of serving live traffic. Nvidia, for its part, is reportedly betting against OpenAI’s push to ban Kimi K3, one more sign that chip supply now shapes alliances as much as raw model quality does.
Nadella’s Bet on Model Independence
Microsoft CEO Satya Nadella laid out the strategy in a lengthy post he titled Frontier Diffusion and Control, which reads less like a product update and more like a doctrine.
“We can now take saturated frontier capabilities and deliver them at scale and at lower cost through models optimized for high-usage products, while continuing to use frontier models for frontier needs,” Nadella wrote. He added that Microsoft is “beginning to route traffic across our first-party surfaces to MAI whenever our models match or outperform frontier alternatives.”
He was careful to note that frontier models from OpenAI and Anthropic remain “part of the orchestration system alongside MAI.” But he also argued that Microsoft’s evaluations should keep improving “even when any given model has been removed,” a principle that only matters if you expect, someday, to remove one.
The subtext lines up with recent history. Reuters reported in April that Microsoft’s once-exclusive license to OpenAI’s technology had been revised into a non-exclusive arrangement. The Information reported last September that Microsoft had begun folding Anthropic models into some of its own products. Microsoft has committed more than $13 billion to OpenAI since 2019. Wednesday’s announcement is the clearest sign yet of what that money is now buying: negotiating power, not dependence.
Developers Like the Idea, Critics Doubt the Execution
On X, developers largely welcomed the pitch behind Nadella’s post.
“I love when people use small models for niche tasks,” wrote developer @mavihsk, responding directly to Nadella. “Why do I have to use the all-knowing model just to change my field in Excel?” Another user, @nabu_lines, put it more bluntly: “cost and performance both improve when you stop overusing the biggest model.”
Designer @designedbyabin was less convinced. “Microsoft is the worst when it comes to listening to user feedback,” they wrote, arguing the company “will lose the AI race because they repeatedly failed to understand user needs.” Another user, @tokenoverflow, turned the model-independence pitch back on Microsoft itself: “i want it keep hill climbing after removing microsoft.”
The skeptics have a fair point on one score. Microsoft’s accept rates, save rates and GPU savings all come from its own internal evaluations, measured against comparisons Microsoft itself chooses to publish. No outside benchmark has replicated any of the headline figures in this announcement.
Microsoft Is Turning the Playbook Into an Azure Product
Nadella isn’t just describing an internal cost-cutting exercise. He positioned the hill-climbing method as “a template for every other AI native, SaaS, or Enterprise company,” and Microsoft is packaging it through Foundry and a new offering called Frontier Tuning, which lets enterprise customers train specialized models against their own proprietary evaluations.
That turns Microsoft’s internal savings into a sales pitch: run AI workloads on Azure, even when the frontier model doing the work comes from somewhere else. Microsoft is also leaning on the claim that its models are trained “on clean, traceable, enterprise-grade data, without distillation from third-party models,” a detail aimed squarely at enterprise buyers and courts increasingly focused on where AI training data comes from.
Microsoft says it’s extending the same approach to Copilot Chat, Outlook and PowerPoint next. Both new models are available now in public preview through Microsoft Foundry and the MAI Playground.
The next quarterly count of patient encounters processed by MAI-Transcribe-1.5 will be a Microsoft number too, produced on Microsoft’s own test set, checked by nobody outside the building.
Frequently Asked Questions
What Is MAI-Transcribe-1.5 Used For Beyond Dragon Copilot?
Per Microsoft’s own Build 2026 keynote, MAI-Transcribe-1.5 is also rolling into Teams, GitHub and Dynamics 365 Contact Center, and it’s available to outside developers through the fastest and most cost-effective transcription model Microsoft offers of any hyperscaler, according to the company.
Has Any Independent Group Verified Microsoft’s Error-Rate Claims?
Not as of this writing. Microsoft has not published a peer-reviewed clinical accuracy study for MAI-Transcribe-1.5. Its public validation history for Dragon Copilot so far consists of Microsoft-commissioned clinician and patient surveys rather than independent clinical trials.
Do Clinicians Still Review AI-Generated Notes Before They’re Filed?
Yes. Microsoft’s own healthcare guidance says clinical staff using Dragon Copilot can pause, edit and validate ambient documentation before it’s filed into the electronic health record, a review step the company positions as a built-in safeguard against errors.
Is Microsoft Cutting Ties With OpenAI Entirely?
No. Nadella has said frontier models from OpenAI and Anthropic remain part of Microsoft’s orchestration system. But Reuters reported in April that Microsoft’s exclusive license to OpenAI’s technology became non-exclusive, and Microsoft is routing a growing share of everyday product traffic to its own MAI models instead.
-
TECHNOLOGY3 years agoHow to Adjust a Bulova Watch Band – An Easy Guide
-
News3 years agoFred Pentland: Athletic Bilbao’s English mentor who changed the essence of Spanish football
-
FINANCE3 years agoTax Planning for Every Season: Guide to Maximizing Your Tax Benefits
-
Education3 years agoAfrican Ministers New Education Plan
-
BUSINESS3 years agoWhat is Entrepreneurial Operating System? A Comprehensive Guide to EOS
-
Education3 years agoInnovate Your Learning Journey with Technology and Enhance Education
-
News3 years agoRussians formally out of World Athletics Championships
-
BUSINESS3 years agoTop 9 Most Expensive American Cities to Rent an Apartment
