groundy
developer tools

GPT-5.6 Is Now Microsoft 365 Copilot's Default: Seat Budgets Can't Assume a Stable Model

OpenAI reportedly set GPT-5.6 as the default for Microsoft 365 Copilot. This rotation breaks static seat budgets by shifting metered costs mid-contract. Version your model in

12 min···4 sources ↓

OpenAI has reportedly said that GPT-5.6 is now the preferred model inside Microsoft 365 Copilot, and if that is true it changes how you should price every seat you have committed to. Two caveats belong up front: the claim comes from a single OpenAI post dated 2026-09-06 that this article could not retrieve and cannot verify, and Microsoft’s own Copilot pages, captured in late August 2026, name no underlying model at all. What follows treats the mechanism, not the announcement, as the durable story.

What does the record actually confirm about GPT-5.6 and Copilot?

Confirmed, in the strict sense: very little. The public pages available at time of writing support three facts and one absence. First, GPT-5.6 exists, Wikipedia’s OpenAI article lists it among OpenAI’s products alongside ChatGPT, which is the only appearance of the model name anywhere in the sources reviewed for this article. That listing says nothing about where GPT-5.6 deploys. Second, Microsoft actively markets Microsoft 365 Copilot, but attributes its behavior to Work IQ, “a workplace intelligence layer that helps Copilot and agents know you, your job and your company”, and to Copilot Cowork, a delegation feature Microsoft’s copy describes as “available, delegate your hardest tasks today.” Third, that same marketing copy names no language model, no GPT version, and no “preferred model” designation of any kind.

The absence is the story. If OpenAI’s post is accurate and GPT-5.6 now sits under Copilot’s preferred-model slot, Microsoft’s buyer-facing surfaces do not say so. Procurement teams reading Microsoft’s Copilot page in the version captured on 2026-08-22, two weeks before the premise post, would find an intelligence-layer brand, a coworker metaphor, and a bundle pitch, cloud storage, security, and Copilot in your favorite apps, “all in one plan”, but nothing that tells them which model their requests execute against. The page may have changed since that capture; this article found no evidence that it has.

This is a familiar pattern in the Microsoft, OpenAI relationship: one party announces, the other’s documentation lags or never arrives. It also means every claim in this article about Copilot’s internals should be read as vendor-asserted until Microsoft publishes routing documentation. The Wikipedia listing confirms the model exists. It does not confirm the deployment. Keep those two sentences separate in your head, because launch coverage this week will not.

What does “preferred model” mean mechanically?

“Preferred model” is a routing designation, not a product name. When a vendor marks a model as preferred inside a hosted assistant, it means the vendor’s serving stack directs default traffic to that model unless something, an admin policy, a capability gate, a fallback rule, overrides the assignment. The user typing into the Copilot pane does not choose; the router does.

For Microsoft 365 Copilot specifically, none of this is documented in the public pages available at time of writing. No learn.microsoft.com page on routing behavior, request metering, or admin model controls turned up, so the mechanics here are inferred from how preferred-model designations work across hosted assistants generally, plus the observable precedent in OpenAI’s own consumer product (covered below). Treat the following as a model of the risk, not a description of Microsoft’s implementation:

  • Assignment is server-side. The vendor decides which model handles a request, and can change that decision without a client update, a contract amendment, or an announcement you are contractually entitled to read.
  • The designation can rotate. “Preferred” is a pointer, and pointers get retargeted. GPT-5.6 today; whatever ships next quarter, later.
  • Cost follows the pointer. If consumption is metered per request and priced per model, then rotating the preferred model rotates the unit price of the same user behavior.

That third point is the one budget owners tend to miss, because seat pricing feels fixed. The seat fee is fixed. What the seat consumes is not.

Why does model rotation move metered spend mid-contract?

Per-seat AI pricing splits into two components: the flat license and the metered consumption that rides on top of it. Model rotation attacks the second component, and it does so mid-contract, when your negotiating position is weakest. Consider what a budget owner actually forecasts. You price N seats at the published per-seat rate, estimate usage from a pilot, apply whatever metered-consumption terms your agreement specifies to the heavy users, and arrive at an annual number. Every term in that calculation assumes the model is a constant. The metering you applied to a Copilot request was calibrated against whatever model served the request when you measured it. If the vendor retargets the preferred-model pointer to a newer, more capable model, one that, on vendor-reported numbers, does more per request, the same user behavior can meter differently. Longer generations, more tool calls per task, agentic features like the delegation workflows Microsoft advertises under Copilot Cowork: each of these raises the consumption attached to a request without the user doing anything differently.

The exposure is asymmetric in the vendor’s favor, and it is worth being blunt about why. A newer preferred model is, by the vendor’s own framing, better. Users who get better answers use the feature more. Usage growth multiplies against whatever metering terms apply to the new model. The budget line you fixed in January is now a function of two variables you do not control: which model serves the request, and how much the improved answers increase request volume. Neither requires a contract change. Both hit the same invoice.

None of this requires malice. Vendors rotate preferred models because newer models are cheaper for them to serve per quality unit, or because the partnership terms changed, or because the old checkpoint is being retired. The reason does not matter to your forecast. The rotation does.

Is vendor-controlled model assignment already the norm?

Yes, and the clearest evidence comes from OpenAI’s own consumer product, not from Copilot. OpenAI’s ChatGPT Android app updated on 2026-09-03, and the app’s Play Store listing carries an older, sharper data point: a paying ChatGPT Plus reviewer wrote on October 11, 2025, “Recently changed and there is no ‘model picker’ for PAID Plus users,” leaving them, in their words, “stuck with 5.”

One review is one data point, and it describes ChatGPT, a consumer surface with different controls and different economics than an enterprise Copilot tenant. It should not be laundered into evidence about Microsoft’s routing. What it establishes is narrower and more useful: OpenAI’s posture as of late 2025 was that model assignment is a vendor decision, and even users paying a monthly subscription did not get an override. Nothing in the pages reviewed for this article shows the picker returning since. The removal of the picker is not a bug in this reading; it is the preferred-model doctrine applied to the consumer app. The vendor prefers a model, so that is the model you get.

Enterprise buyers have historically assumed they sit outside this posture, that contracts, admin consoles, and compliance requirements carve out an exception. Maybe they do for Copilot. The public pages reviewed here do not say. What the consumer precedent tells you is that the default, absent a negotiated control, is vendor assignment. If your Copilot agreement does not explicitly grant you model selection or pinning rights, assume you do not have them, because the party writing the serving code has demonstrated what it does with unilateral control.

What should you demand from Microsoft before seat renewals?

The procurement ask reduces to four questions, each mapping to a control that either exists in writing or does not exist at all. Verbal assurances from an account team are not controls.

1. Routing documentation. Which model serves which Copilot surface, and what does “preferred” override? Microsoft’s current marketing copy attributes Copilot to Work IQ, which is positioning language, not a routing architecture. If the answer to “which model executed this request” is a brand name, you cannot audit anything. Ask for the documentation page, not the slide.

2. Multiplier schedules per model. If premium requests are metered, the multiplier attached to each model is a pricing term. Get the current schedule if one exists, get the change-notification terms attached to it, and get in writing what happens to in-flight commitments when the schedule changes. A schedule that can be revised with 30 days’ notice is a variable, and it belongs in your model as one.

3. Admin model pinning. Can a tenant administrator pin Copilot traffic to a specific model version, at what granularity (tenant, group, surface), and for how long is a pinned version supported? This is the single most valuable control in the list, because it converts model rotation from a vendor decision into a customer decision. If pinning exists, price it. If pinning exists only on a higher tier, that price difference is the premium Microsoft charges for budget stability, and it is worth computing explicitly.

4. Change notification for preferred-model rotation. Not documentation after the fact, notification before it. If OpenAI’s 2026-09-06 post is how customers learn about model changes, customers are learning from the wrong vendor.

A vendor that cannot answer these four questions in writing is telling you something about where the control actually lives. Believe it.

How should your cost model version the model as a line item?

Stop treating the underlying model as a constant in per-seat forecasts and start treating it as a versioned variable with a rotation scenario attached. Concretely, the structure that survives a preferred-model rotation looks like this:

  • Base seats at the contracted rate, the only truly fixed term.
  • Metered consumption per seat, estimated from pilot usage, with any metering or multiplier terms cited by version and access date, the same way you would cite a dependency in a lockfile.
  • A named-model assumption: which model your consumption estimate was measured against, stated explicitly in the budget document. When the preferred model rotates, this is the term that invalidates, and you want it findable.
  • A rotation scenario: re-run the consumption estimate under the assumption that the preferred model moves to a newer, higher-consumption default mid-contract. If agentic delegation features (the Copilot Cowork pattern) drive more requests per task, that growth belongs in the scenario, not in a footnote.
  • A pinning line: the cost of the admin control or tier upgrade that buys model stability, compared against the rotation scenario’s downside. This turns pinning from an IT preference into a priced insurance decision.

The discipline here is the same one engineering teams already apply to software dependencies. Nobody ships a production system with an unpinned latest tag and then acts surprised when the build changes. Seat budgets have been running on the equivalent of a floating tag, because until recently there was one plausible model and the fiction of stability held. The GPT-5.6 announcement, if accurate, is the moment the fiction stops holding. Version the model. Pin it where the contract allows. Re-run the forecast when the vendor rotates, not when the invoice arrives.

When is pinning an older model the cheaper path?

Pinning is cheaper whenever the consumption delta between model versions exceeds the cost of the pinning control, and that condition is more common than vendor framing suggests, for a specific reason: most Copilot usage is not frontier-model work. Summarizing a long thread, drafting a reply, reformatting a table, the workloads that dominate a typical tenant’s request distribution are well inside the capability envelope of the previous model generation. The marginal quality a newer preferred model adds to those tasks is real but small. The marginal consumption, if the newer model generates longer, calls more tools, or meters at a higher rate, may not be small.

This is the calculation launch coverage will not do for you, because it requires your usage distribution rather than the vendor’s benchmark. A tenant where most Copilot requests are commodity office tasks has a different answer than a tenant where Copilot Cowork delegation runs long multi-step workflows. The first tenant should price pinning aggressively and probably take it. The second may genuinely benefit from riding the preferred-model pointer and should instead spend its negotiation effort on multiplier schedules and notification terms.

There is also a verification burden worth naming. Pinning only works if you can confirm the pin holds. If Microsoft provides no way to observe which model served a request, and no such observability appears in the public pages available at time of writing, then a pinned tenant is trusting the same routing layer it was trying to escape. Demand the observability alongside the pin. An unauditable control is a settings page, not a control.

What is the verdict, and what would change it?

Version the model inside your per-seat cost models, treat mid-contract rotation of the preferred model as a live budget risk, and negotiate routing documentation, multiplier schedules, and pinning controls before your next annual commit. The pattern underneath is durable and does not depend on the GPT-5.6 claim being true: vendors control model assignment even for paying users, as OpenAI’s removal of the ChatGPT model picker in late 2025 demonstrated on its own consumer surface, and “preferred model” designations are server-side pointers that rotate without contract changes. Budgets built on a stable-model assumption will be wrong; the only question is whether they are wrong in your favor.

Now the limitation, and it is a large one. Every Copilot-specific mechanism in this article rests on a single OpenAI vendor post that this article could not retrieve, and on inference from how hosted assistants generally work. Microsoft’s own pages, captured in late August 2026, attribute Copilot to Work IQ and Copilot Cowork and name no model, no multiplier schedule, and no admin control. One Microsoft documentation page, describing actual routing behavior, actual premium-request metering for Copilot, or an actual pinning toggle, could invalidate the specifics here in an afternoon. The budgeting framework survives that page. The GPT-5.6 narrative may not. Build the framework, and check for the documentation before your renewal, not after the invoice that tells you it existed.

Frequently Asked Questions

Does the removal of the ChatGPT model picker apply to Microsoft 365 Copilot tenants?

No, the picker removal is specific to OpenAI’s consumer ChatGPT app. Microsoft 365 Copilot is a distinct enterprise surface where model assignment is governed by Microsoft’s routing logic and Work IQ layer, not OpenAI’s consumer UI controls. The consumer precedent indicates a vendor preference for unilateral assignment, but it does not confirm that Copilot lacks admin pinning features, which remain undocumented in public sources.

How does the ‘Work IQ’ branding affect the ability to audit model routing?

Work IQ is a marketing label for an intelligence layer, not a technical routing specification. Because Microsoft’s public pages attribute Copilot behavior to Work IQ without naming the underlying LLM, procurement teams cannot verify which model version executes a request. This opacity prevents auditors from confirming whether a ‘preferred model’ rotation actually changed the inference engine, making it impossible to validate cost models against actual serving behavior.

What specific documentation should be requested to verify Copilot’s metering terms?

Request the specific learn.microsoft.com page detailing premium-request multipliers and per-model pricing schedules. Since no such documentation was found in the public record as of September 2026, the absence of this page means metering terms are currently contractual placeholders. Without this document, any budget forecast assuming static per-request costs is speculative, as the vendor could alter multiplier rates for newer models without public notice.

Why is the three-day gap between the ChatGPT app update and the GPT-5.6 announcement significant?

The gap is likely coincidental, as the app update on 2026-09-03 predates the vendor post naming GPT-5.6 as Copilot’s preferred model. The app update relates to OpenAI’s consumer product, while the Copilot claim concerns Microsoft’s enterprise stack. Treating the timing as causal would conflate two separate product lines and vendors, leading to incorrect assumptions about how consumer app changes impact enterprise Copilot routing or pricing.

sources · 4 cited

  1. OpenAIen.wikipedia.orgcommunityaccessed 2026-09-06
  2. Microsoft 365 Copilot | AI Productivity Tools for Workmicrosoft.comvendoraccessed 2026-09-06
  3. ChatGPT - Apps on Google Playplay.google.comvendoraccessed 2026-09-06