The Cache Got Cheaper Than the Model
Return to Part 0: Table of ContentsPrevious Article: Part 36, Permanent Access, Except For You
Two frontier labs published new prices in the same week, and the number that travelled fastest was the wrong one.
On 21 September 2026, xAI released Grok 4.7. It has 2.1 trillion parameters, forty per cent more than Grok 4.6's 1.5 trillion, trained in part on supplemental SpaceX engineering data. The price held: two dollars and six dollars per million input and output tokens, the same rate xAI already charged.[1][2] The next day, Anthropic cut pricing on Claude Opus 5.5. The figure that circulated by evening was "forty per cent cheaper." It is not on Anthropic's own pricing table. Nothing on that table says forty.[3]
What the table actually says: base input pricing fell twenty per cent, five dollars to four. Output fell twenty per cent too, twenty-five dollars to twenty. Cached context, the material a long-running AI agent rereads every time it checks its own working memory, fell sixty per cent, fifty cents to twenty cents per million tokens.[3] Had Anthropic simply carried the twenty per cent base cut through to caching at the old rate, a cache read on the new four-dollar input price would cost forty cents. It costs twenty. That second cut targets caching specifically, not the base price it is drawn from.
A vendor that discounts everything by the same amount is being generous. A vendor that discounts one mechanism twice as steeply as everything else is showing its hand. It reveals which shape of AI system it expects to win.
What a cache multiplier actually prices
Every model call is priced by the token, and for a single chat question that is a small, forgettable number. An autonomous agent working through a codebase, a contract, or months of a customer's support history behaves differently. It rereads the same large block of context on every step. The model carries no memory between calls unless that material is handed back to it each time. Left unpriced, that repetition would make long-running agents expensive to operate for no analytical reason.
Prompt caching is the industry's answer: store the context once, charge a fraction of the base rate for every subsequent read. Across Anthropic's model line, and across most of the industry, the standard fraction has been one tenth of the base input price. Write the cache close to full price, read it back for roughly a dime on the dollar.[3]
Claude Opus 5.5 broke that ratio. Anthropic's own pricing documentation, fetched directly from the live table, lists base input at four dollars per million tokens, five-minute cache writes at five dollars, one-hour cache writes at eight dollars, output at twenty dollars, and cache hits at twenty cents. That is a multiplier of five hundredths rather than the standard tenth. A footnote on the same table states it without qualification: cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.[3] The launch announcement carries the identical figures.[4] Two independently fetched Anthropic pages agree.
Claude Opus 5, for comparison, sits at the standard rate: five dollars input, twenty-five output, fifty cents cache read, exactly ten per cent of the five-dollar base.[3] Line the two models up and the twenty per cent cuts to input and output track each other precisely. The cache-read cut does not. Sixty per cent is three times the size of the twenty per cent cuts sitting either side of it on the same table. The only way to get there is to move the multiplier itself, not the base price it is a fraction of.
That is a specific claim about the future, whether Anthropic states it that plainly or not. Short chat sessions barely touch the cache mechanism. The tokens involved are too few to matter. Long-lived agents live inside it, rereading a large persistent context many times across a single working session. Pricing that mechanism at half its old rate, while leaving the mechanisms a chat product depends on at a flat twenty per cent discount, is a bet. It is a bet on which workload shape Anthropic wants more of.
Put a number to it. A compliance-review agent working through a two-hundred-page policy corpus rereads that corpus on every clause it checks, not once at the start. Under the old 0.1x multiplier, a session that touches the cached corpus four hundred times across a working day carries a meaningfully different bill than the same session under 0.05x. The base input and output rates, meanwhile, moved by only twenty per cent each. A single customer-support reply, by contrast, barely enters the cache at all. There is no large persistent document behind it, so the multiplier that just halved is close to irrelevant to what that call actually costs. Two workloads, running on the same model, at the same headline price cut, experience two different economics. That gap, not the headline twenty per cent, is what a total-cost model has to price.
The correction that makes the argument stronger, not weaker
Three weeks earlier, on 1 September 2026, Anthropic released Claude Fable 5.1 and its restricted-access sibling Claude Mythos 5.1. Their cache-hit price on the same live table: twenty-five cents per million tokens against a ten-dollar base, a multiplier of 0.025, four times below the standard tenth and steeper again than Opus 5.5's own 0.05.[3] Independent reporting on the Fable 5.1 launch corroborates the figure: prompt-cache reads falling from a dollar to twenty-five cents, a four-times cut, at launch on 1 September.[5] The footnote naming Opus 5.5 as an exception sits on the same page as the footnote naming Fable 5.1 and Mythos 5.1 as a separate, steeper exception. Anthropic has now moved the cache multiplier below its own standard on three models, across two release dates, inside a single month. The deeper cut went to the earlier, less prominently marketed release.
No public statement ties those two pricing decisions together as one announced policy, and it would overstate the evidence to claim otherwise. What can be said is narrower and more useful. This is the same mechanism repriced twice in three weeks, on two different tiers of the same company's current line-up. It reads less like an isolated marketing decision and more like a standing view about where persistent context sits in Anthropic's own architecture. I made a version of this argument in July, when a multi-vendor price war showed that once the model itself is a commodity, durable value moves into the identity and audit layer wrapped around it. What has changed since then is the grain. Anthropic is now pricing one specific piece of that surrounding layer, the part that holds an agent's working memory between calls, differently from everything else it sells.
What this means for the architecture decision in front of you
If a vendor is willing to price one mechanism at half or a quarter of its standard rate, that is worth examining at design time. Ask whether your architecture is rewarded for using that mechanism, or penalised for avoiding it. Three questions belong in that conversation.
First, does the workload actually behave like the thing being discounted? A persistent-context agent, one that keeps a project's codebase, a client's case history, or a policy corpus resident across many calls, earns real money back from a 0.05x or 0.025x cache multiplier. A workload built from short, independent calls barely touches the mechanism. It gains almost nothing from the headline cut, whatever a summary slide says about the model being cheaper overall.
Second, what does the discount cost you in lock-in? A cache-hit price this favourable is an incentive to keep context resident on one vendor's infrastructure for as long as possible. Moving it, rebuilding it on a competitor's cache, or re-architecting around a shorter context window all forfeit the saving. That is a reasonable commercial decision for Anthropic to make. It is a decision an enterprise architect should recognise as a lock-in mechanism before signing anything, not after the second invoice arrives.
Third, is the comparison in front of your board honest about what changed? The twenty per cent base cut and the sixty per cent cache cut are two different decisions, not one blended discount. Treating them as a single "forty per cent cheaper" line in a board paper will misprice any total-cost-of-ownership model built on it. Run the arithmetic on your own token mix, not the headline figure.
None of this needs to stay theoretical to be useful. Pull last month's token bill, split it into base input, output and cache reads, and recalculate each line at the new rate before assuming the vendor's own percentage applies to your workload. An organisation running mostly short, stateless calls will find the real saving closer to twenty per cent than sixty. An organisation running a small number of long-lived agents against large, stable context will find the opposite. The cache line moves the most, and it is also the line most exposed if the multiplier moves again at the next release. Knowing which of those two profiles describes your usage, rather than assuming, is what this exercise is for.
For a New Zealand organisation weighing an agentic pilot against a production commitment, none of this is abstract. Most agentic architecture decisions made this year assumed cache pricing was a stable, minor input, a rounding error next to the base token rate. That assumption no longer holds. The multiplier is now the single line item most likely to move at the next pricing revision, and it moved twice in three weeks on one vendor's books alone. Treat it as an architectural variable to be reviewed at each pricing revision, not a constant fixed at procurement time. Vendor diversity is not only a resilience question for the model itself. It is increasingly a resilience question for the specific mechanism an agent's memory depends on.
The contrast case, and what to watch next
Not every vendor made the same bet this week. xAI's Grok 4.7 arrived a day before Opus 5.5 with forty per cent more parameters and no price change at all. It held the same two dollars and six dollars per million input and output tokens xAI already charged, absorbing a substantial capability increase into the existing rate card rather than passing any of it through as a discount on any mechanism.[1][2] Reporting on the release briefly disputed the parameter count: one outlet's headline said 210 billion against the 2.1 trillion nearly every other source, including that outlet's own article body, reported. The 2.1 trillion figure survives a check of the source against itself.[6] Two vendors, one week, two different pricing postures. One held the line on a larger model. The other restructured how it charges for a specific piece of infrastructure, twice, with no public sign either watched the other.
Anthropic has already cut the multiplier twice in three weeks. The open question is whether a second vendor follows: whether a discounted cache multiplier becomes something buyers come to expect from every frontier lab, the way discounted batch processing already has, or whether it stays an Anthropic-specific bet on persistent-context agents that the rest of the market declines to match. That is a checkable question, with a real answer sitting on a small number of public pricing pages.
The forty per cent figure that circulated this week was wrong, but understandably so. It is easier to remember than "twenty per cent here, sixty per cent there, and twice in a month besides." The real numbers took longer to find and said more.
The open-source dimension sits one layer down from the price cut. The mechanism Anthropic just repriced, holding a large context resident and serving repeated reads of it cheaply, is not proprietary in the way the label suggests. Open-source inference servers such as vLLM, built at Berkeley and now maintained as an Apache-licensed project under the PyTorch Foundation, implement the same prefix-caching approach. Any organisation willing to run its own inference stack can use it, on hardware it controls, at a cost it sets rather than one a vendor's footnote can move twice in three weeks. That does not make self-hosted inference free, and frontier-scale weights still come from closed labs such as Anthropic and xAI. What it does do is put the caching mechanism itself, not only the model, inside the sovereignty question. An organisation running its own cache is not exposed to the next multiplier change, whoever moves it.
The sovereignty implication follows the same logic one step further. If a vendor prices its infrastructure to reward whoever keeps a large, persistent context resident on its own servers, the jurisdiction that server sits in stops being a procurement footnote. It becomes a standing question about where an organisation's working memory actually lives. Sovereign AI initiatives such as the United Kingdom's AI Security Institute and Japan's AI Safety Institute have already begun treating compute geography, not only model access, as a governance variable. A cache-pricing regime that rewards long residency on one vendor's infrastructure sharpens that question rather than settling it. For a defence or public-sector agency evaluating an agentic deployment, the multiplier attached to cached context is a small technical fact with a direct line to a much larger one. Whose infrastructure, in which jurisdiction, is quietly becoming the default home for an institution's accumulated context.
What has your organisation's own agent architecture assumed about the price of memory, and has anyone actually gone back to the rate card to check?
If your organisation is moving AI agents from pilot to production and nobody outside the vendor has inspected the control plane, message me and I will send the scope and the fixed fee for an independent review.
The views expressed in this article are entirely my own, informed by more than 30 years of professional experience in architecture, security, and technology leadership in New Zealand. I write as director of Te Pono Limited; the views are personal and do not represent the position of any client, any government agency, or the New Zealand government. My commentary on legislation and policy is analytical, drawing on publicly available sources and my professional expertise in architecture, security, and AI governance, and it is politically neutral.
Andreas Hamberger is a New Zealand leader in Architecture & Security and Associate Member of the Institute of Directors. The Hamberger Report: Generative AI 2026 provides enterprise leaders with evidence-based analysis of the AI landscape. Through Te Pono he provides independent reviews of agentic AI control planes for organisations moving from pilot to production; contact andreas@thehambergerreport.com for the scope and fixed fee.
This article was produced with AI assistance under my direction. Research, drafting and images pass through a pipeline I built and govern: automated gates for source verification, forbidden language and political neutrality, and my own review before anything is published. The tools include Claude, Gemini and Openart. The frameworks, arguments and editorial judgements are mine and are the same discipline I apply to the AI systems I audit for clients. AI accelerated the work; the thinking, and the responsibility for it, are mine.
[1] xAI. "API Pricing." 2026. https://docs.x.ai/developers/pricing
[2] Decrypt. "xAI Launches Grok 4.7. It's Bigger, But Late to the AI Frontier Party." 21 September 2026. https://decrypt.co/378824/xai-launches-grok-4-7
[3] Anthropic. "Pricing." Claude Platform Docs. 2026. https://platform.claude.com/docs/en/about-claude/pricing
[4] Anthropic. "Introducing Claude Opus 5.5." 22 September 2026. https://www.anthropic.com/claude-opus-5-5
[5] MacRumors. "Anthropic Launches Claude Fable 5.1 With Lower Costs and Fewer False Positives." 1 September 2026. https://www.macrumors.com/2026/09/01/anthropic-claude-fable-5-1/
[6] KuCoin. "xAI Launches Grok 4.7 with 210 Billion Parameters, Incorporates SpaceX Data Training." 21 September 2026. https://www.kucoin.com/news/flash/xai-launches-grok-4-7-with-210b-parameters-adds-spacex-data-training

