Two Flagship AI Models, 42 Days Apart. The Real Changes Never Reached a Benchmark Chart. | Edition 308
Edition 308 — Anthropic shipped Opus 4.8 just 42 days after 4.7. Same price, same context, same cutoff. What changed was the plumbing.
Anthropic released Claude Opus 4.7 on 16 April 2026, and Claude Opus 4.8 on 28 May 2026. Two flagship models, 42 days apart.
Coverage of that second launch put the gap at 41 days. It is 42 — April has 30 days, so 14 remain after the 16th, and May contributes 28. Nobody was harmed by the missing day, but it is a fair preview of how carefully the rest of the numbers around these two releases have been handled.
Because here is the more interesting problem. The single most-quoted benchmark figure for Opus 4.8 — a precise percentage on SWE-bench Verified, repeated confidently across write-ups — does not appear anywhere on Anthropic’s announcement page for Opus 4.8. Not in the table, not in a footnote. We checked the page directly for the term. It is not there.
So this edition does something less exciting than a scoreboard and considerably more useful. It reads the documents nobody writes headlines about — Anthropic’s dated release notes, its pricing tables, its token-count tables — and works out what actually changed between these two models. The answer is almost nothing you could put on a benchmark chart, and quite a lot that shows up on an invoice.
What Anthropic actually shipped, and when
Anthropic’s platform release notes are dated and public, which makes them the most reliable record of this period. Read in order, they tell a stranger story than either announcement did.
- 16 April 2026 — Claude Opus 4.7 launches, at what the release note calls “the same $5 / $25 per MTok pricing as Opus 4.6.” It adds a new
xhigheffort level, task budgets in beta, and high-resolution image input up to 2,576 pixels on the long edge. - 12 May 2026 — fast mode, already a research preview on Opus 4.6, is extended to Opus 4.7.
- 28 May 2026 — Claude Opus 4.8 launches, 42 days after 4.7, with “the same set of tools and platform features as Claude Opus 4.7.”
- 9 June 2026 — Claude Fable 5 launches, 12 days later, described in Anthropic’s documentation as its most capable widely released model.
- 25 June 2026 — fast mode is deprecated for Opus 4.7, with removal scheduled for 24 July.
- 24 July 2026 — Claude Opus 5 launches, described as “a step-change improvement over Claude Opus 4.8,” at the same $5 / $25 pricing. The same day, fast mode is removed from Opus 4.7 entirely.
Read that sequence again with one detail in mind: fast mode arrived on Opus 4.7 on 12 May and was gone by 24 July. A feature lifetime of 73 days — on a model that was already 16 days away from being superseded when the feature landed.
Where the benchmark numbers actually come from
Anthropic’s Opus 4.8 announcement does publish a comparison table. It also carries four benchmark claims in its running text: a Super-Agent benchmark, CursorBench, a Legal Agent Benchmark, and a score of 84% on Online-Mind2Web.
Every one of those four appears inside a testimonial from a named customer engineer or executive — not in Anthropic’s own description of what it measured.
That is not a scandal. Customer testimonials are a normal part of a launch page, and the people quoted are presumably reporting exactly what they saw in their own systems. But a number that reached you through one customer’s evaluation harness is a different class of evidence from a number a lab measured and published with its methodology attached, and the two belong in separate mental buckets. The first tells you what happened to somebody. The second tells you what to expect.
Then there is the figure that is not there at all. Search for Opus 4.8 and you will be handed a precise SWE-bench Verified percentage. We are not going to reprint it here, even with a warning attached, because reprinting a number is how it becomes true. It is not on Anthropic’s page. If Anthropic publishes a SWE-bench result for this model, it will be quotable then.
What actually changed between 4.7 and 4.8
Two documented differences, both from Anthropic’s own materials, and both of them overheads that got smaller. That is a shorter list than the coverage of this launch suggests, and the reason it is shorter turns out to be the most useful thing on this page.
1. Tool calls got substantially cheaper to set up
Every API request that includes tools carries a hidden system prompt that Anthropic adds automatically — and Anthropic publishes its exact size, per model, in the pricing documentation. On Opus 4.7 it was 675 tokens with a tool choice of auto or none, and 804 tokens with any or tool. On Opus 4.8 the same two figures are 290 and 410.
Take the common case, a tool choice of auto: that is 385 fewer input tokens on every tool-using request, before your own prompt is counted — a 57% reduction in that overhead. (The any/tool pair falls by 394 tokens, or 49%.) For a low-volume application it is invisible. For an agent making tool calls in a loop, it compounds; the calculator above does the arithmetic at Anthropic’s published $5 per million input tokens. The bash tool definition, for what it is worth, costs the same 325 additional tokens on both models.
2. The prompt-caching floor dropped
Prompt caching only applies above a minimum prompt length, and Anthropic’s prompt-caching documentation publishes that minimum per model: 1,024 tokens on Opus 4.8, down from 2,048 tokens on Opus 4.7. Prompts that were too short to cache on 4.7 became cacheable on 4.8 — and a cache hit costs 10% of the standard input price.
The two differences that turned out not to be differences
Two further changes are widely attributed to Opus 4.8. Neither survives a reading of the documentation.
Adaptive thinking is not a 4.8 feature. Anthropic’s thinking documentation lists it for Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6 and Claude Sonnet 4.6 alike: on every one of them, thinking is off until you set the thinking type to adaptive, which lets Claude decide when and how deeply to think. Whatever else separates these two models, this does not.
Effort defaulting to high is not a change either. Anthropic’s effort documentation gives the API default as high in its Opus 4.7 section and in its Opus 4.8 section, in the same four words. The Opus 4.8 announcement concedes the continuity itself: on coding tasks, it says, that effort level spends a similar number of tokens as Opus 4.7’s default, but with better performance. Migrating an application between these two models does not quietly change what effort costs you.
What is genuinely new is a control rather than a default: a new control alongside the model selector lets users choose how much effort Claude puts into a response, on claude.ai and in Cowork. That is a setting a person can now reach, not a hazard hiding in an API default.
Which is this edition’s own lesson, pointed back at it
We began this comparison expecting four differences and finished with two. The other two arrived exactly the way that SWE-bench percentage arrives: through coverage, stated confidently, sourced to nobody. One of them was written down here as a direct quotation — a sentence that appears in no Anthropic document we can find, and that was still sitting in this draft when the fact-check reached it.
So the honest version is duller than the one we set out to report, and rather more useful. By Anthropic’s own description these two models carry the same tools and platform features, the same prices, the same context window and the same knowledge cutoff. What separates them is two quiet overhead reductions that no launch post led with — and several of the differences you have read about between them did not happen at all.
The thing that didn’t change at 4.8 — but did at 4.7
Anthropic’s pricing documentation carries a note that is easy to scroll past: Claude 4.7 and later models use a newer tokenizer, one that “produces approximately 30% more tokens for the same text.” That figure is the headline average. The Opus 4.7 announcement gives the per-content-type range as roughly 1.0–1.35× — the two are the same change described at different resolutions, not competing numbers.
This matters for reading the price table correctly. Opus 4.6, 4.7 and 4.8 all list at $5 and $25 per million tokens — a perfectly flat line across three releases. But the document that cost you a million tokens on Opus 4.6 costs roughly 1.3 million on 4.7. The price per token held steady; the number of tokens did not.
The important detail for anyone comparing these two specific models: that step happened at 4.7, not 4.8. It is not a 4.7-to-4.8 difference, and anyone upgrading between those two was already paying it. It is, however, a very good example of why “the price didn’t change” and “the bill didn’t change” are different sentences — and why a benchmark chart would never have told you either way.
| What the coverage said | What the record shows |
|---|---|
| Opus 4.8 shipped 41 days after Opus 4.7 | 42 days. Anthropic’s platform release notes date Opus 4.7 to 16 April 2026 and Opus 4.8 to 28 May 2026. April has 30 days, leaving 14 after the 16th, plus 28 in May. |
| Opus 4.8 scores a specific percentage on SWE-bench Verified | The term “SWE-bench” does not appear anywhere on Anthropic’s Opus 4.8 announcement page. The figure circulates on aggregator write-ups only, so it is not reprinted here. |
| Opus 4.8’s fast mode is three times cheaper than Opus 4.7’s | Anthropic’s pricing page lists fast mode for Claude Opus 5 and Claude Opus 4.8 only, at $10 / $50 per million tokens, and states it is not available on Opus 4.7. Fast mode was removed from Opus 4.7 on 24 July 2026. |
| Opus 4.8 was a major capability jump over 4.7 | Anthropic’s own release note describes Opus 4.8 as having “the same set of tools and platform features as Claude Opus 4.7.” The documented differences are two token overheads — the tool-use system prompt and the prompt-caching floor — not platform capability. |
What to actually do with this
If you run anything on the Claude API, three of these are worth ten minutes of your afternoon.
Check whether you set effort explicitly. If you do not, and you are on a model where it defaults to high, you are running a spending default you never chose. Setting it deliberately — in either direction — beats inheriting it.
Revisit your prompt-caching assumptions. If you once concluded your prompts were too short to be worth caching, that conclusion was reached against an older, higher minimum.
Measure tokens, not documents. Any cost model you built before Opus 4.7 understates what the same workload costs now, even though the headline price never moved.
And then the broader point, which is the one worth carrying out of here. Both of these models are now filed under “Legacy models” in Anthropic’s own documentation. Opus 4.7 held the flagship slot for 42 days. Twelve days after Opus 4.8 arrived, Claude Fable 5 landed; fifty-seven days after it, Claude Opus 5.
In a release cycle moving this fast, a benchmark chart is a photograph of a single afternoon — and half the ones you will be shown were taken by somebody else. The release notes, the pricing tables and the token-count tables are dull, dated, and unglamorous. They are also the part that tells you what it costs to live there.