← All posts AI & Automation 18 min read

Claude vs GPT vs Gemini for Business in 2026: Compare the Retirement Clocks, Not the Benchmarks

Model comparisons expire faster than the integrations they inform. The current Claude, GPT-5.6 and Gemini model IDs and prices as of August 2026, and the vendor lifecycle policies — notice periods, published retirement dates, and the platform split — that decide what a shutdown actually costs you.

Title card for the guide comparing Claude, GPT and Gemini model retirement policies
Quick Summary

Claude vs GPT vs Gemini for Business — Compare the Retirement Clocks, Not the Benchmarks

  • As of August 2026 the current families are Claude (claude-opus-5, claude-sonnet-5, claude-haiku-4-5), OpenAI’s GPT-5.6 series (gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna), and Gemini 3.x (gemini-3.6-flash, gemini-3.5-flash, gemini-3.5-flash-lite). Every model named in a comparison written a year ago is either retired or superseded.
  • Capability rankings are the wrong axis for a build decision, because they expire faster than the integration they inform. Gemini 2.0 Flash shut down on 1 June 2026 and Claude 4 was retired on 15 June 2026. A comparison published that month was recommending models that had already stopped answering requests.
  • The durable difference between the three vendors is what they promise before switching a model off. Anthropic commits to at least 60 days’ notice and also publishes an earliest-retirement date for models still in service. OpenAI commits to at least 6 months for generally available models but as little as 2 weeks for preview models. Google publishes dated shutdown tables but states no guaranteed minimum notice.
  • Choosing a cloud changes your retirement dates, by months. Claude Sonnet 4 was retired on Anthropic’s own API on 15 June 2026; AWS lists the same model on Amazon Bedrock with an end-of-life date of 14 October 2026. One model ID, two published dates, four months apart.
  • Per-token prices are not comparable across model generations. Anthropic documents that Claude 4.7 and later use a tokenizer producing roughly 30% more tokens for the same text, so a cheaper headline rate can be a more expensive invoice.
  • The decision that survives: pin exact model IDs, keep the provider call behind one interface, and diary the published retirement dates. Pick the tier by context window and caching shape, not by benchmark position.
60 days
Anthropic’s stated minimum notice before retiring a publicly released model, per its model deprecations page
2 weeks
The notice OpenAI states a preview model may get, against at least 6 months for a generally available one
4 months
Gap between Claude Sonnet 4’s retirement on Anthropic’s API (15 Jun 2026) and its end-of-life on Amazon Bedrock (14 Oct 2026)
~30%
Extra tokens the same text produces on Claude 4.7 and later, per Anthropic’s pricing note — headline rates are not like-for-like

Every model comparison is a photograph of a moving object. The useful question for a business is not which model scores highest this quarter; it is which vendor will give you enough warning when the model you built on is switched off. That answer is published, it is different at all three vendors, and it does not appear in any of the comparison articles currently ranking for this question.

This guide gives the current model IDs and prices as of August 2026, then the part that outlives them: the three vendors’ published lifecycle policies side by side, the platform split that quietly changes your dates, and the token-accounting change that makes headline prices misleading. It ends with a portability checklist you can work through this week. Model selection is one decision in a larger stack — the AI for commerce teams guide hub covers the rest.

Contents

What the current models actually are, as of August 2026

Three families are current. Anthropic lists claude-opus-5, claude-sonnet-5 and claude-haiku-4-5-20251001 as active, alongside claude-fable-5 as its most capable widely released model. OpenAI’s current flagship series is GPT-5.6, in three tiers: gpt-5.6-sol, gpt-5.6-terra and gpt-5.6-luna. Google lists gemini-3.6-flash, gemini-3.5-flash, gemini-3.5-flash-lite and gemini-3.1-flash-lite as generally available.

Two spec facts matter more than any ranking. Anthropic documents a 1M-token context window on Claude Fable 5, Opus 5 and Sonnet 5, and 200k on Haiku 4.5. OpenAI documents 1.05M tokens across the GPT-5.6 tiers. Those numbers decide which jobs are possible at all; benchmark positions decide which are marginally better.

Names carry no version information, which is the first practical trap. gpt-4o is still listed and still sold at $2.50 per million input tokens, but the gpt-4o-2024-05-13 snapshot underneath it was deprecated on 22 April 2026 with a shutdown date of 23 October 2026. A dated snapshot ID and an undated family name behave differently under a deprecation, and only one of them is safe to hard-code.

Why a model comparison expires before the integration does

An integration written against a model outlives the comparison that chose it, usually by years. The comparison rots first, and it rots on a schedule the vendors publish in advance.

The dates are specific. Google lists gemini-2.0-flash, released 5 February 2025, with a shutdown date of 1 June 2026, and gemini-2.0-flash-lite, released 25 February 2025, on the same shutdown date. Anthropic records claude-sonnet-4-20250514 and claude-opus-4-20250514 as deprecated on 14 April 2026 and retired on 15 June 2026. Any article published in June 2026 recommending “Gemini 2.0” or “Claude 4” was recommending endpoints that had already stopped answering, or were days from it.

This is not a criticism of the vendors, who published every one of those dates ahead of time. It is a statement about what a comparison can and cannot carry. A capability verdict is true for a quarter. A lifecycle policy is true for the life of the contract, and it is the thing that determines whether a shutdown is a scheduled afternoon of work or an outage. Treat the model choice as reversible and the vendor choice as the commitment.

What each vendor promises before it switches a model off

All three publish a deprecation policy, and they are not equivalent. The table below is assembled from the three vendors’ own documentation pages, checked 11 August 2026. Read it as the contract you are signing, because it is the part of the choice that does not change next quarter.

Lifecycle question Anthropic (Claude API) OpenAI Google (Gemini API)
Committed minimum notice, generally available model At least 60 days At least 6 months None stated
Committed minimum notice, preview or specialised model Not separately stated At least 3 months for specialised variants; as little as 2 weeks for preview models None stated
Earliest retirement date published for models still current Yes — a “not sooner than” date per active model No — models appear once deprecated Yes — a dated shutdown table, stated as earliest possible dates
Published lifecycle vocabulary Active, Legacy, Deprecated, Retired Legacy, deprecated, shut down Stable (GA), preview, shut down
Partner clouds follow the same dates No — Bedrock and Google Cloud set their own Not addressed on the deprecations page Not addressed for partner models

Verdict: OpenAI gives the longest guaranteed reaction window on generally available models and the shortest on preview ones, so the safety of an OpenAI build depends entirely on which tier you pinned. Anthropic gives the shortest stated guarantee of the two vendors that commit to one, but it is the only vendor of the three that publishes both a guaranteed notice period and a forward-looking earliest-retirement date. Google publishes dates without committing to a notice period, so its planning horizon is good and its floor is undefined.

Notice period and planning horizon are two different guarantees

A notice period and a planning horizon answer different questions, and most teams conflate them. The notice period tells you how long you get to react once the announcement lands. The planning horizon tells you, today, the earliest date a model you are already running could disappear. You need both, and no vendor’s marketing frames it this way.

Anthropic’s page states the reaction guarantee directly: it notifies customers with active deployments “at least 60 days” before retirement for publicly released models. Its status table simultaneously supplies the planning horizon, listing tentative retirement dates such as “Not sooner than July 24, 2027” against models that are currently active.

Google inverts the pair. Its deprecations page publishes dated shutdowns well ahead of time but states that those dates “indicate the earliest possible dates on which a model might be retired” and that the exact date will be communicated “with advance notice” — no floor attached. You can plan against the table; you cannot size a migration sprint against the guarantee, because there is not one.

OpenAI supplies the strongest floor and the weakest horizon: at least 6 months for generally available models, but the deprecations page lists models once they are already deprecated rather than publishing an earliest date for those still current. In practice that means an OpenAI build is safe to plan in half-year increments and impossible to plan beyond them.

Guaranteed migration windows published by three model vendorsThree timeline lanes from deprecation announcement to shutdown. OpenAI guarantees at least six months for generally available models but as little as two weeks for preview models. Anthropic guarantees at least sixty days. Google states no guaranteed minimum, publishing earliest possible shutdown dates instead.Guaranteed window between announcement and shutdownDrawn to scale against a six-month axis. Solid: a committed minimum. Dashed: no committed minimum.announcement6 months laterOpenAIgenerally availableat least 6 monthsOpenAIpreview modelsas little as 2 weeksAnthropicpublicly releasedat least 60 daysGoogleGemini APIno committed minimum — dated table published instead

The same model can be retired on one platform and running on another

Where you buy the model changes when it disappears. Anthropic’s deprecation page states this explicitly: its dates apply to Anthropic-operated platforms, while “partner-operated platforms (Amazon Bedrock and Google Cloud) set their own retirement schedules, so a model’s lifecycle status and dates can differ.”

AWS says the same thing from the other side. Its Bedrock model lifecycle page notes that dates there “are specific to Amazon Bedrock and may differ from dates published by model providers (such as Anthropic or Cohere)”, and that for Bedrock usage “only the dates on this page apply”. Both vendors are telling you the same thing: there is no single retirement date for a model, only a retirement date per platform.

The gap is measurable. Anthropic records claude-sonnet-4-20250514 as retired on its own API on 15 June 2026. AWS lists anthropic.claude-sonnet-4-20250514-v1:0 on Bedrock with a Legacy date of 14 April 2026 and an end-of-life date of 14 October 2026 — four months of additional life for an identical model, published by both parties, reconciled by neither.

The platform guarantees differ in structure too, not just in dates. Bedrock commits that a model “will remain on Amazon Bedrock for at least 12 months before the EOL date” and that it “will be in the Legacy state for at least 6 months before the EOL date” — a longer floor than the model provider’s own 60 days. It also adds a stage that has no first-party equivalent: after at least three months in Legacy, a model enters public extended access, where AWS states you “should expect higher pricing, which will be set by the model provider”. A retirement is not always a cliff. Sometimes it is a price rise you have to notice on an invoice.

This breaks the standard ecosystem advice, which says to pick the model your cloud already carries and stop thinking about it. Picking the cloud is not a neutral convenience choice; it silently reassigns the retirement date that governs your integration, in either direction. A team on a partner cloud can outlive the first-party retirement by a third of a year, and a team that migrates from a partner cloud to the first-party API can inherit an earlier shutdown than the one it planned against. Read the lifecycle table for the platform you actually bill through, not the one the model was announced on.

Why per-token prices are not comparable across model generations

A price per million tokens only means something if a token is the same size on both sides of the comparison. Anthropic documents that this stopped being true within its own lineup: Claude 4.7 and later models “use a newer tokenizer” that “produces approximately 30% more tokens for the same text”, with the exact increase depending on content and workload shape.

Work the arithmetic on a fixed document rather than on a rate card. A body of text that costs 1,000,000 tokens under the older tokenizer costs roughly 1,300,000 under the newer one. At an identical headline rate of $5 per million input tokens, the same document moves from $5.00 to about $6.50 — a 30% increase in the invoice with no change in the advertised price. Comparisons that rank vendors by rate card alone will get this backwards whenever a tokenizer changes underneath one of them.

The correction is to price the workload, not the rate. Take a representative sample of your real inputs — 50 support tickets, 50 product descriptions, whatever the job actually is — send them through each candidate model’s token counting endpoint, and multiply actual token counts by current rates. That is the only comparison that survives a tokenizer change, and it takes an afternoon.

What the three families cost today

Published list prices as of 11 August 2026, per million tokens, from each vendor’s own pricing page. These are the most volatile numbers in this guide; re-check them before you commit a budget.

Tier Claude GPT-5.6 Gemini
Highest-priced current model, input / output Fable 5 — $10 / $50
Opus 5 — $5 / $25
Sol — $5 / $30 3.5 Flash — $1.50 / $9
Mid-priced, input / output Sonnet 5 — $2 / $10 to 31 Aug 2026, then $3 / $15 Terra — $2 / $12 3.6 Flash — $1.50 / $7.50
Lowest-priced, input / output Haiku 4.5 — $1 / $5 Luna — $0.20 / $1.20 3.5 Flash-Lite — $0.30 / $2.50
Cache read 0.1x base input 0.1x base input Separate per-token fee, from $0.025
Cache write 1.25x base input for 5 minutes; 2x for 1 hour 1.25x base input Included in the caching fee
Cache storage held over time No time-based charge No time-based charge $1.00 per 1M tokens per hour

Verdict: the low tiers are not within rounding distance of each other — GPT-5.6 Luna’s input rate is a fifth of Claude Haiku 4.5’s — so high-volume classification work is worth pricing properly rather than defaulting. The row that changes architecture, though, is the last one.

Google is the only one of the three that bills cached context by time as well as by token. Deriving from its published rate of $1.00 per million tokens per hour: a 200,000-token cached system prompt held continuously costs $0.20 per hour, or about $4.80 per day, whether or not a single request arrives. Anthropic and OpenAI charge a one-off write multiplier instead, so an idle cache costs nothing. For a support assistant with a large fixed prompt and bursty overnight traffic, that difference decides the design — on Google you cache around traffic windows, on the other two you cache and forget.

Matching a commerce workload to a model tier

Match the tier to the shape of the job, using documented specs rather than benchmark position. Four workload shapes cover most commerce and ERP work.

  • High-volume, low-judgement classification — order tagging, refund-reason coding, product categorisation. Lowest tier, batched. Anthropic’s Batch API takes 50% off both input and output; price the batch rate, not the standard one.
  • Large fixed prompt, many small requests — a support drafting assistant carrying a policy document on every call. Caching-dominated, so the caching shape in the table above matters more than the base rate.
  • Whole-document reasoning — reconciling a long statement against ERP records, or reviewing a full contract. Context-window bound: Claude Fable 5, Opus 5 and Sonnet 5 document 1M tokens and the GPT-5.6 tiers document 1.05M, which decides feasibility before cost does.
  • Anything writing to a system of record — creating sales orders, adjusting inventory, issuing refunds. Model choice is secondary here; the governance limits of the receiving system dominate, as covered in the guide to calling AI APIs from SuiteScript.

What deliberately is not in that list is a capability ranking. Benchmark scores for these families move monthly, are measured on tasks that rarely resemble order reconciliation, and cannot be verified from a vendor’s own documentation. Context window, price, batch discount and caching behaviour are all documented and stable enough to design against. Build on the documented numbers and treat capability as something you measure on your own inputs.

The portability checklist

Portability is what converts a retirement announcement from an incident into a ticket. Work down this list once, and the next shutdown costs an afternoon.

  • Pin an exact model ID in configuration, never a floating family name, so a silent substitution cannot change behaviour underneath you.
  • Keep the model ID in environment configuration rather than in code, so switching it is a deploy and not a release.
  • Route every model call through one internal interface — one function, one module — so the provider-specific request shape exists in exactly one file.
  • Record which platform issues the credential for each workload, because the retirement date follows the platform and not the model.
  • Diary the published retirement or earliest-shutdown date for every model ID you run, and set the reminder before the vendor’s notice period would start.
  • Keep a stored set of representative inputs and expected outputs, so a replacement model can be evaluated the day it is announced rather than the week it is forced.
  • Subscribe the on-call address, not an individual’s inbox, to each vendor’s deprecation notifications.
  • Confirm no prompt depends on an undocumented quirk of one model, since that is the dependency a migration discovers last.

How to audit your own exposure this week

Two hours of work establishes whether a model retirement would be routine or disruptive. The audit is mechanical and needs no vendor contact.

Start by listing every model ID actually in production, taken from usage exports rather than from the codebase — Anthropic’s Console, for instance, exports usage broken down by API key and model, which surfaces the forgotten script that no repository search finds. Then, for each ID, open the vendor’s deprecation page and record its current state and its published retirement or earliest-shutdown date. Finally, note which platform bills each workload, and check that platform’s own lifecycle table rather than the vendor’s headline one.

The output is a short table: model ID, platform, lifecycle state, published date, owner. Anything with a date inside twelve months gets a migration ticket now; anything with no published date at all is the higher risk, because it is the case where the vendor’s committed notice period is the only warning you will get. Teams that keep this table treat model retirements as maintenance. Teams that do not, discover them from a 404 in production.

Get the working checklists

The runbooks and decision checklists from these guides, as printable PDFs — free in the SoftXone guide library.

Browse the guide library →

When this framing does not apply

Lifecycle discipline is overhead, and three situations do not justify it. If the only use is a person typing into a chat interface, there is no integration to break — the vendor swaps the model underneath and the work continues. If the deployment is a genuine prototype with a decision date inside a quarter, pin nothing and pick on capability, because the project will be rewritten or cancelled before any retirement lands.

The third exception is the interesting one. If the model is reached through a tool someone else built, the retirement risk is that vendor’s to manage and yours to inherit without notice, which is a different exposure than the one described here and is examined in the analysis of what sits behind AI-powered SaaS tools.

Everywhere else — anything on a schedule, anything writing to an ERP, anything a customer touches — the retirement clock is the governing constraint. The same reasoning applies one layer down to the protocols these models speak, where a revision can remove a mechanism entirely rather than just a model, as happened to session handling in the Model Context Protocol specification. If you want this designed once and handed over rather than assembled in-house, that is the integration work we do.

Sources & Further Reading

References

  1. Anthropic — Model deprecationsLifecycle vocabulary, the 60-day notice commitment, the partner-platform statement, and the dated retirement table for Claude models.
  2. Anthropic — Models overviewCurrent Claude model IDs, context windows, and max output tokens.
  3. Anthropic — PricingPer-token rates, prompt caching multipliers, the Batch API discount, and the tokenizer note on Claude 4.7 and later.
  4. OpenAI — DeprecationsThe minimum notice periods for generally available, specialised and preview models, and the definitions of legacy, deprecated and shut down.
  5. OpenAI — ModelsCurrent GPT-5.6 model IDs and context window sizes.
  6. OpenAI — PricingPer-token input, cached input and output rates for the GPT-5.6 tiers and gpt-4o.
  7. AWS — Amazon Bedrock model lifecycleBedrock’s Active/Legacy/EOL states, the 12-month and 6-month floors, the public extended access pricing stage, and the dated EOL table including Claude Sonnet 4.
  8. Google — Gemini API model deprecationsDated shutdown tables including gemini-2.0-flash, and the statement that listed dates are earliest possible dates.
  9. Google — Gemini modelsCurrent Gemini model IDs and their stable or preview status.
  10. Google — Gemini API pricingPer-token rates and the context caching storage charge per million tokens per hour.

Frequently asked questions

Can I keep using a model after its retirement date?

Generally no — after the shutdown date, requests to that model fail. Two documented exceptions are worth knowing. AWS states that after a Bedrock end-of-life date the model is unavailable in all regions “unless there is a private arrangement between you and the provider for continued access”. Bedrock also runs a public extended access period before end-of-life, during which active users keep access at pricing the model provider sets, typically higher than before. Neither exception is automatic, and both require you to have noticed well in advance.

What happens to a fine-tuned model when the base model becomes legacy?

On Amazon Bedrock, customisation stops but existing deployments keep running. Once a base model enters the Legacy state you cannot start new fine-tuning jobs against it and cannot create new Provisioned Throughput endpoints, though deployments created before the transition continue until the end-of-life date. Two access rules catch teams out: new customers cannot begin using a Legacy model at all, and existing customers may lose access after 15 days of inactivity, so an occasional batch job can find its model gone before the published date.

Can an API change break my integration without any model being retired?

Yes, and deprecated request parameters are the usual cause. Anthropic documents that temperature, top_p and top_k are deprecated on Claude Opus 4.7 and later, and that setting any of them to a non-default value returns a 400 error rather than being silently ignored. Code that ran unchanged for a year against an older model can fail on its first request to a newer one, with nothing retired and no shutdown date involved. Audit request parameters whenever you move a model ID, not only the model name.

How do I find every model my organisation is actually calling?

Use billing data rather than a code search, because scheduled jobs and one-off scripts rarely appear in the main repository. Anthropic’s Console exports a CSV of usage broken down by API key and model from its Usage page, and equivalent per-model usage breakdowns exist in the other vendors’ consoles. Reconcile that export against your deployment inventory. Anything appearing in the export but not in the inventory is the exposure you did not know you had, and it is usually the thing that breaks.

Is the model with the lowest price per million tokens the cheapest to run?

Not reliably, because two published mechanisms break the comparison. Tokenizers differ between model generations: Anthropic documents roughly 30% more tokens for the same text on Claude 4.7 and later, so an identical headline rate produces a larger bill. Caching is billed differently too. Google charges context caching storage at $1.00 per million tokens per hour, a time-based charge that accrues even on an idle cache, while Anthropic and OpenAI apply a one-off write multiplier and no storage charge.

Related guides

Discussion

Leave a Reply

Your email address will not be published. Required fields are marked *


Ship it

Need this in your stack?

We build, integrate, and ship — no calls, just delivery.

Start a project →