01 AI Consulting 02 Software Development 03 About 04 Blog
DE EN
Arrange a call
← All posts

AI Consulting

Two Models in One Day: What September 22 Reveals About the AI Market

On September 22, 2026, two vendors released new models within one to two hours of each other: Anthropic Claude Opus 5.5, OpenAI the tiers GPT-6 Sol and GPT-6 Luna. Which house went first is contradictorily documented; the shared day is certain. Neither vendor has stated any intent behind it.

More interesting than the question of who won is a measurement published the same day. The independent analysis service Epoch AI calculated across five benchmarks how fast the price for a given level of performance falls: by roughly 47 percent per quarter, about thirteenfold per year since 2023. That is the real news of this day — and it means that any business case twelve months old is working with prices from a different era.

At a glance

  • What happened: On 22.09.2026 Claude Opus 5.5 appeared alongside GPT-6 Sol and Luna — but not flagship against flagship: OpenAI showed mid-sized and small tiers at roughly halved prices, while its top model GPT-6 Astra dates from September 3.
  • What follows: Epoch AI measures 47 percent price decay per quarter at equal performance — faster than the historical cost curves for electricity, compute and batteries.
  • The catch: The independent predeployment evaluation of Opus 5.5 calls the progress “incremental” — and discloses that the vendor under review held an editorial right of review over the report.

What happened that day

DateAnthropicOpenAI
July 24, 2026Claude Opus 5—
September 1, 2026Claude Fable 5.1 and Mythos 5.1—
September 3/4, 2026—GPT-6 Astra (flagship)
September 22, 2026Claude Opus 5.5GPT-6 Sol + GPT-6 Luna

The reading “two top models in a duel” does not survive scrutiny: OpenAI showed no new top end, but two cheaper tiers below its own flagship — which remained GPT-6 Astra from September 3. Anthropic, conversely, presented its flagship, 19 days after Astra and barely two months after Opus 5. What the two announcements share is not the performance peak but the price.

The number that explains the day

Epoch’s measurement is the most robust figure of the quarter because it does not come from a vendor and because it is internally consistent: 47 percent per quarter compounds to roughly a factor of 13 over four quarters. Part of that is hardware — every dollar invested in AI chips delivers, according to Epoch, about 49 percent more compute per year. The decay in model usage is therefore faster than the curve beneath it; what goes beyond it comes from better methods — my reading, not Epoch’s measurement.

On individual prices I deliberately stay silent. My sources contradict each other: the list prices imply a drop of about one fifth for Opus 5.5, while the vendor itself speaks of 40 percent lower running costs — depending on the reference basis, both can be true. And the figures for the GPT-6 tiers come from a single source that a second round of checking could not find again. The list price is not the price you pay anyway: cache reads cost only 2.5 to 10 percent of the input price, and batch processing gives a 50 percent discount.

Three objections to the simple reading

First: “independently evaluated” carries less than it sounds. METR, one of the few bodies that evaluate models before release, tested Opus 5.5 over ten working days on five tasks. The verdict: better than its more expensive predecessor, but explicitly an “incremental improvement.” More remarkable than the verdict is what METR writes about its own limits: alignment properties were not evaluated — and the vendor under review held an editorial right of review over the published text. An external evaluation whose wording the evaluated party helps edit is not a certification. Anyone justifying due-diligence obligations with “externally evaluated” should know what that sentence is worth — the same pattern as with the safety narrative of a vendor.

Second: prices do not only fall. An introductory price is a sales measure, not a basis for calculation: the one for Gemini 3.8 Flash runs only until 31.12.2026 and doubles on January 1, 2027, and DeepSeek actually raised the prices for its V4-Pro model in August and started billing by peak and off-peak hours for the first time.

Third: switching is development work, not a line of configuration. In the current Claude generation, temperature, top_p and top_k have been removed — divergent values return an error, and almost every older LLM project sets them. On top of that, the deprecations: Opus 4.1 has been switched off since August 5, Sonnet 4.5 follows no earlier than September 29, 2026, Haiku 4.5 from October 15.

A horizontal bone-colored measuring bar with a finely notched edge lies against deep ink black on top of a solid, filled bone surface. A piece is missing from the middle of the bar; the removed segment floats slightly offset above it and is drawn only as a hollow outline with the same notches. At the right-hand cut edge, a single short vertical vermilion mark runs down into the surface.

The price decay is measured. Whoever measures it is not always independent: in the predeployment evaluation, the evaluated party held an editorial right of review over the report — a piece of the yardstick has been cut out and set aside.

Who is actually working inside your Copilot?

Since September 25, three days after the double launch, Microsoft has been routing its Copilot automatically: software selects, by accuracy, speed and cost, which model answers — from the GPT family, the Claude family or Microsoft’s own MAI line. Whoever does not know which model is working does not know who processes the data; the Auftragsverarbeitungsvertrag (data processing agreement), the Verarbeitungsverzeichnis (record of processing activities) and the third-country assessment all wobble at the same point. Demand the information on which models the router may use, and notification when the list changes. (Dated secondary source; I have not verified Microsoft’s own announcement.)

What this means for your business

First: get old rejections out of the archive. If a project failed on usage costs in 2025, that calculation is not slightly wrong today but from a different order of magnitude. Recalculate the three or four most expensive cases.

Second: do not hard-wire yourself. Only those who can switch benefit from the price decay, and the binding rarely forms at the model but at the context around it — prompts, memory, tool integrations. How expensive this context lock-in becomes is underestimated by most projects.

Third: make model transparency a requirement. Record in contracts and in your AI policy which models are permitted in which tool, and who informs you about a change.

Conclusion

September 22 was not a duel but a pricing round — and the most interesting number of the day came from neither participant. If the price for a given level of performance really falls by about half per quarter, then “which model is the best?” is the wrong question. The right one is: which projects pay off now that did not pay off a year ago?

One caveat belongs with this: not everything falls, introductory prices expire, and the evaluation that classifies the progress is itself not entirely independent. Whoever plans on that basis plans with ranges. If you want to know which of your discarded AI projects hold up under today’s conditions, let’s talk — I do that arithmetic as a developer and read the contractual and data protection consequences as a business lawyer.

FAQ

What was released on September 22, 2026?

Anthropic presented Claude Opus 5.5, OpenAI the tiers GPT-6 Sol and GPT-6 Luna — on the same day, one to two hours apart. Flagship against flagship it was not: OpenAI showed mid-sized and small tiers at roughly halved prices, while its top model GPT-6 Astra dates from September 3, 2026. Which house published first is contradictorily documented.

How fast are prices for AI capability actually falling?

On September 22, 2026, Epoch AI measured across five benchmarks that the price for a given level of performance falls by roughly 47 percent per quarter — about thirteenfold per year since 2023. Both figures describe the same movement: 47 percent per quarter compounds to roughly a factor of 13 over four quarters.

Why does this article name no prices per million tokens?

Because my sources contradict each other: the list prices imply a drop of about one fifth for Claude Opus 5.5, while the vendor itself speaks of 40 percent lower running costs. The figures for the GPT-6 tiers come from a single source that a second round of checking could not find again. What holds up is the direction and the rate of the decay, not the individual prices.

What does this mean for model selection in a company?

Old rejections on cost grounds belong back on the agenda, not in the archive. But only those who can switch benefit from the decay, and switching is development work: in the current Claude generation the parameters temperature, top_p and top_k have been removed. Anyone buying a tool with automatic model selection additionally needs the assurance of which models the router may use.


Sources — as of 26/09/2026

Note: openai.com was not reachable for this research (HTTP 403); the OpenAI facts rest on the developer changelog and third-party sources. The absolute prices of the GPT-6 tiers were single-sourced and not confirmed in a second pass — which is why this article names no token prices.

This article is general information and not legal advice in an individual case. The market moves fast; as of September 26, 2026 — please check prices and deprecation deadlines at the primary source before making decisions.

Leon Lotz

Leon Lotz

Leon Lotz is a business lawyer and founder of MusketierSoftware. He combines legal depth with real software craft.

AI-assisted, editorially reviewed and under editorial responsibility. AI transparency