Not all model upgrades are upgrades

Source: Microsoft for Developers By Waldek Mastykarz
Image: Microsoft for Developers

Summary

A new model drops with lower per-token pricing and better benchmarks. You switch. A week later someone asks why the agent is burning 12x more tokens on the same task while producing worse output. We ran 150 agent tasks across 15 scenarios on two models, Claude Sonnet 4.6 and Claude Sonnet 5, using GitHub Copilot […] The post Not all model upgrades are upgrades appeared first on Microsoft for Developers.

Read this on Microsoft for Developers
Opens in a new tab. Subvolts summarizes and links; the full piece belongs to Microsoft for Developers.
Where the other five stand

Related: OpenAI: Advancing the price-performance frontier with GPT-5.6 · Google: Introducing Gemini Robotics 2 · Anthropic: Claude Fable 5 - Full 319 page Breakdown · Meta: This AI-Powered Wheelchair Could Change How People Move · xAI: Introducing Grok 4.5: Fast, affordable intelligence

Hype check
2/5Worth a look

Rated low: routine. Worth knowing, not worth rearranging your day for.

Published
Source
Microsoft for Developers (devblogs.microsoft.com)
Author
Waldek Mastykarz
Company
Microsoft · Official · Developer
Summary by
Subvolts, using an extract from the source (how we work). Spotted a mistake? Tell us.

Questions people ask

Where can I read the full story?
On Microsoft for Developers. The "Read this on Microsoft for Developers" link above opens the original in a new tab. Subvolts publishes a summary and analysis, never the full piece.
What does this mean for Copilot?
A new model drops with lower per-token pricing and better benchmarks. You switch. A week later someone asks why the…

More from Microsoft for Developers 9 more

Page generated Sep 3, 2026. Summaries are Subvolts' own; the story belongs to Microsoft for Developers.