AI Transformation19 Aug 202619 min read

How to Reduce Claude API Costs Without Losing Quality

Prompt caching, the Batch API, the effort setting, compaction and model routing: the design choices that decide your Claude API bill, updated for Opus 5.5.

How to Reduce Claude API Costs Without Losing Quality

Last updated: 7 Oct 2026.

Most Claude API bills are set the week the system is designed, not the week it goes live. By the time a finance director asks why the invoice is three times the pilot estimate, the decisions that caused it, which model handles which step, whether context gets cached, how hard every request is allowed to think, are already baked into code nobody wants to touch.

That's the uncomfortable bit for a finance-literate buyer: there is no procurement lever that fixes a badly designed system after the fact. API cost is an architecture decision, not a contract term: the same workload can cost several times more or less depending on how it is built, with the vendor's price list unchanged either way. What follows is the set of levers that actually move the number, updated for the Claude Opus 5.5 release on 22 Sep 2026, for the person holding the budget and the engineers holding the build.

Why can't we just negotiate a better rate?

Because the API is metered per token, and the token count is a function of what your system asks the model to do, not who's asking. Two companies on identical published rates can see very different bills for the same outcome: one routes every request through its most expensive model at full reasoning depth, while the other tiers its work, caches what repeats, and batches what can wait. Anthropic's pricing page is the reference for the rates themselves. This post is about everything that sits between those rates and your invoice.

This is also a separate question from what a Claude subscription costs. If you are choosing between Team and Enterprise seats, or working out what a Claude plan costs a UK business, our guide to Claude Enterprise pricing for UK buyers answers that. Subscription pricing licenses people to use Claude. API pricing meters what your own software asks the model to compute, at whatever volume your product drives. Confusing the two is how a discovery call ends with a subscription quote for a workload that was always going to run through the API on a different cost shape.

Two billing facts UK buyers ask about. Anthropic bills API usage in US dollars, so your accounts team will see exchange-rate movement on top of usage movement, and any pounds figure in a budget is an estimate. Tax follows the billing address you register in the Claude Console, and Anthropic's help centre describes an optional field for a tax or VAT ID where the address qualifies. Anthropic's pages do not state a UK rate, so treat VAT as a question for your accountant rather than something the price list answers.

Which model should each step run on?

Anthropic sells more than one size of model, and the largest is not the default answer. Our rule on client builds: start with the mid-tier model for anything that isn't trivially simple, and move up to the largest, or down to the smallest, only when an evaluation shows the move pays for itself. Anthropic's own guide to optimising for cost and intelligence makes the same point from the other direction: price lists are written per token, but you pay for completed tasks, so the comparison that matters is cost per completed task, including retries.

"Pays for itself" is doing real work there. It means running the same representative inputs through both models, scoring the outputs against a standard you'd defend to a client, and comparing the token cost of the difference, not an engineer's hunch that the flagship model "feels more reliable" (usually true, rarely worth the multiple). Defaulting to the cheapest model to protect margin is the same mistake in reverse: the saving is lost several times over in retries once it fails on harder inputs. Anthropic positions its smallest current model, Claude Haiku 5.5 (released 7 Oct 2026), for high-volume and narrowly scoped work, including subagent tasks, and says Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding, which is a fair summary of where we land too. Haiku 5.5 is also priced by prompt length: prompts over 100,000 tokens pay five times the rate, so it suits short, frequent calls far better than long documents. We compare it with OpenAI's equivalent in Claude Haiku 5.5 or GPT-6 Luna.

Two things changed on model choice in September 2026. Claude Opus 5.5 launched with a list price a fifth lower than Claude Opus 5 on both input and output tokens, so the gap between the top tier and the middle narrowed. And Anthropic notes that its models from Claude 4.7 onwards use a newer tokeniser that produces roughly 30 percent more tokens for the same text, which means a per-token price comparison across model generations understates the newer model's cost on the same workload. Compare on cost per task, measured on your own inputs, not on the rate card. Our living comparison of current AI models for UK business keeps the model landscape current, and if you are already running Opus 5, our guide to moving from Claude Opus 5 to Opus 5.5 covers what breaks and what changes when you switch.

The other half of routing is recognising that not every step needs a model call at all. Formatting a date, validating a postcode, checking a number against a threshold: none of that needs a large language model, and routing it to one means paying model-token prices for logic a line of code would handle for free. Ask whether each pipeline step has a right answer code can compute before assuming it has to go through Claude. We go into where that boundary should sit in our guide to Claude AI agents for business.

How does prompt caching cut the bill?

Prompt caching is the highest-leverage lever for any system sending the model a large, stable block of context on every request: a lengthy system prompt, a product catalogue, a set of tool definitions. If the model has already processed an identical block of text recently, it reuses that processing rather than reading the block again at full price. In Anthropic's own measurements, published in the cost guide linked above, caching was the largest single lever by a wide margin, cutting agent-loop cost by a factor of between 2.7 and 5.3 on Anthropic's benchmarks. Those are Anthropic's numbers on Anthropic's workloads, so treat them as an indication of the ceiling, not a forecast for yours.

The pricing structure, from Anthropic's prompt caching documentation and its pricing page, is simple enough to reason about. Writing a block to the cache costs a quarter more than a normal input token for a five-minute cache, or double for a one-hour cache. Reading it back costs a tenth of the normal input rate on most models. On Claude Opus 5.5, and on Claude Sonnet 5.5 since 7 Oct 2026, a cache read costs a twentieth of the input rate, half the usual multiplier, which makes the case for caching on those models stronger still. Anthropic's arithmetic is that the five-minute cache pays for itself after a single read and the one-hour cache after two. Opus 5.5 will only cache a prompt of at least 512 tokens, so very short system prompts gain nothing from it.

The catch is that the match has to be exact. Caching works on a prefix: everything up to the point you mark cacheable must be identical to what was cached before. Change a single character anywhere in that block, an unintentional timestamp, a reordered field, a genuinely updated instruction, and the cached prefix is invalidated. Nothing errors: it just runs at full price, silently, which is how systems that stitch a current date into the "stable" part of the prompt never see a hit and never know it. Expiry does the same thing more quietly: the first request after a gap longer than the cache lifetime misses and re-bills the whole context. Anthropic now offers an automatic mode, where a single top-level cache_control field lets the API manage cache breakpoints as a conversation grows, alongside explicit breakpoints on individual content blocks. Either way, getting the benefit means separating what is genuinely constant (put it first, mark it cacheable) from what varies, then checking the cache_read_input_tokens figure the API returns on every response to confirm the cache is actually being hit.

Anthropic's pricing page carries a worked example that shows caching's limits as well as its value: for a one-hour agent session where 40,000 of 50,000 input tokens come from cache, the total cost of the session falls by roughly a quarter. Not more, because output tokens, which caching never touches, are most of that particular bill. Treat headline case-study percentages with the same care. Anthropic's customer page for Notion leads with a 90 percent cost reduction from caching, but the same figure appears in Anthropic's own caching launch post as a generic worked example (a cached 100,000-token document), and Notion's quoted statement carries no number. Before you build a business case on someone else's percentage, ask which workload it was measured on.

What workloads suit the Batch API?

Not everything needs an answer in the next two seconds. Overnight enrichment of a dataset, bulk classification of historical records, drafting hundreds of similar documents, evaluation runs: anything tolerating a delay of minutes to hours rather than an instant reply is a candidate for the Message Batches API instead of the standard real-time endpoint. Anthropic charges half price on both input and output tokens for batched requests. Most batches complete within an hour, but the guarantee is only that results arrive within 24 hours, and requests still unprocessed at that point expire. Results stay downloadable for 29 days.

The batch discount stacks with caching. Because a batch can take longer than five minutes to run, Anthropic suggests using the one-hour cache when batched requests share a large common prefix, and it is honest that cache hits inside a batch are best-effort because the requests run concurrently. For a nightly job with a stable system prompt, the combination is the cheapest way to run Claude that exists. Anthropic's cost guide calls batching the second-largest free lever after caching for unattended agent work.

The judgement call is honest classification of urgency. A support agent answering a live chat cannot wait for a batch window. A nightly job re-scoring every lead in the CRM can, and paying full real-time price for that job buys latency nobody asked for. Businesses running a serious volume of Claude API calls usually have at least one workload sitting in the wrong lane. We break down where the money goes in our piece on what AI agents cost in the UK.

How does the effort setting change what a request costs?

The reasoning the model does before it answers is billed as output tokens, at the full output rate, and you are billed for the whole thinking process whether or not your application displays any of it. That makes reasoning depth a spend setting that feels like a quality setting, and on current models the control for it is the effort parameter, with five levels from low to max. Effort shapes every output token, including tool calls and the text of the answer, so lower effort also means fewer and terser tool calls. If your team is still thinking in terms of a thinking token budget, that mental model is out of date: on the newest models a manual budget is rejected, and effort is the lever.

Two facts about Claude Opus 5.5 matter here. Its default effort is medium, where Opus 5, Fable 5.1 and Sonnet 5.5 default to high (Haiku 5.5 also defaults to medium), so an integration that moves to Opus 5.5 without setting effort explicitly runs one level lower than it did, cheaper and shallower, without anyone deciding that. And thinking cannot be switched off on Opus 5.5: a request that tries returns an error at every effort level. Teams that used "thinking disabled" as their cheap setting on earlier models now lower effort instead. Anthropic's advice, which we would give anyway, is to run an effort sweep on your own evaluation set rather than carry a setting over from a previous model.

The mistake we see is one effort level applied across an entire pipeline because it worked best on the hardest case in testing. Anthropic's own measurements on its knowledge-work benchmarks found medium matching the default's accuracy at roughly 70 to 87 percent of its cost, and low giving up one to three points for a third to a half off. Vendor benchmarks again, but the shape is the point: where a workflow mixes genuinely difficult judgement calls with a larger volume of routine steps, set effort per task, not once globally. One interaction to know about before you do: changing the top-level effort value between requests invalidates the prompt cache, because the setting is rendered into the prompt. Vary effort across workloads rather than within a cached conversation, or on the models that support it use the per-message effort change, currently in beta, which preserves the cache. Finally, max_tokens is the hard ceiling on thinking plus answer combined, the one setting that bounds a single request's spend absolutely, so set it deliberately.

What do compaction and context editing do for long-running agents?

In a long agent conversation the input side of the bill grows every turn, because every turn re-sends the whole history. Two Anthropic features address that directly, and neither was available when this post was first written.

Compaction replaces the older turns of a conversation with a summary that Claude writes on the server, so you write no summarisation code of your own and the active context stays small. It is in beta in two forms. Compaction on demand, requested with the compact-2026-09-04 beta header and a top-level compaction parameter, lets your application decide when to summarise, can run in the background while work continues, and can keep the most recent turns word for word so only the older history is condensed. Compaction at a token threshold hands the decision to the API: you set a trigger and it summarises inside the ordinary request that crosses it. Anthropic recommends the on-demand form wherever it is available. One cost note from Anthropic's own documentation for Claude Code, which uses the same idea: compacting a large context is itself a large request, because the model has to read everything it is about to summarise. Compaction saves money over the turns that follow, not on the turn that runs it.

Context editing is the blunter tool. Once a conversation passes a threshold you set (Anthropic's default is 100,000 input tokens), the API clears the oldest tool results, keeping the most recent few, and replaces each with a short placeholder. File contents and search results the model has already acted on are exactly the sort of bulk that accumulates in an agent loop and never needs re-reading. The interaction to plan for is with caching: clearing content invalidates the cached prefix, so Anthropic's guidance is to clear enough at once to make that worthwhile, using the clear_at_least setting, rather than trimming a little on every turn. For how context size affects quality as well as cost, see our explainer on Claude's context window.

Do multi-agent systems really cost more, or is that overstated?

It's understated, if anything. Anthropic's write-up of the multi-agent research system it built states that agents typically use about four times the tokens of a chat interaction, and multi-agent systems, where a lead agent coordinates several subagents in parallel, about fifteen times. In the same evaluation, a multi-agent setup with a stronger lead model coordinating smaller subagents outperformed a single agent on the same task by over ninety percent.

Both numbers are true at once, and the design decision is holding them next to each other rather than picking the one that suits the pitch. A fifteen-times multiplier is defensible when the task is high-value, genuinely parallelisable, and would otherwise exceed a single context window: substantial research synthesis, competitive analysis across many sources. It is indefensible on a routine task a single well-scoped agent, or no agent at all, would handle for a fraction of the cost. Where a multi-agent design does earn its place, the levers above still apply inside it: Anthropic's effort documentation names subagents as the typical case for low effort, and tool definitions shared across subagents are a natural cache prefix. The step we insist on before building a multi-agent architecture is asking, in pounds, whether the task in front of you is one where that multiplier pays for itself. How to test an AI agent covers the evaluation discipline that answers it.

What about Priority Tier, fast mode and data residency?

Three items on Anthropic's pricing pages that look like cost levers and are not. Priority Tier, the paid capacity commitment that once bought guaranteed throughput, is no longer available for purchase: existing commitments run to their contract end, and Anthropic's docs point anyone needing guaranteed capacity to its sales team. If a plan or a proposal you are reading still recommends it, that advice is out of date. Fast mode, a research preview that returns Opus output faster, is priced at a premium, double the standard rate on Opus 5.5, and is not available through the Batch API. It buys speed at extra cost. And pinning inference to the United States with the inference_geo setting adds a tenth to every token category on Claude 4.6 and later models, so only set it when a data-residency requirement, not habit, demands it.

Why do averages mislead on API cost, and what should we track instead?

Because API cost distributions are skewed, not evenly spread. A small share of requests, unusually long documents, an unexpected multi-step tool-use chain, a retry loop, account for a disproportionate share of total spend. A monthly average cost per request can look respectable while a tail of outliers quietly drives the bill up, and nobody notices until the invoice arrives.

The fix is instrumentation from day one, not after the first surprising bill. Log token usage per request, tagged by workflow step and by model, from the first production deployment, so when the invoice moves you can point at exactly which step or model tier caused it. The API gives you the raw material on every response: separate counts for cached and uncached input, cache writes, output, and (on current models) the share of output that was thinking. It's the fastest way to catch a cache silently failing to hit, a step routed to the wrong model tier, an effort setting nobody chose, or a batch-shaped job still running in real time. Anthropic also exposes an organisation-level usage and cost API for finance reporting, and a free token-counting endpoint for estimating a prompt's size before sending it. Instrumenting this properly as part of a build, rather than bolting it on afterwards, is core to how we scope work through our Claude implementation service.

The cheapest lever is often upstream of the model: contact data quality and AI running costs explains how duplicate and incomplete records inflate the bill before a single design decision is made.

Frequently asked questions

Does a bigger model always give a better result, so isn't tiering down a quality risk?

Not necessarily, and treating it as an automatic trade-off is itself a cost mistake in reverse. On many steps, formatting, extraction, routine drafting, a mid-tier or smaller model matches the larger model's output at a fraction of the cost, and the only way to know is to run both through an evaluation against the same inputs and compare cost per completed task, retries included. The larger model earns its higher cost on the steps where an evaluation shows it actually changes the outcome.

What changed about cost when Claude Opus 5.5 launched?

Three things. Its list price is a fifth lower than Claude Opus 5 on both input and output tokens. Cache reads cost a twentieth of the input rate rather than the usual tenth, so caching is worth more on this model than on most Claude models (Sonnet 5.5 has matched it since 7 Oct 2026). And its default effort is medium rather than high, so an integration that upgrades without setting effort explicitly will run cheaper and shallower than before, which is worth deciding rather than discovering.

Can prompt caching and batch processing be used together?

Yes, and Anthropic confirms the two discounts stack. Because a batch can take longer than five minutes to run, use the one-hour cache when batched requests share a large common prefix, and expect cache hits inside a batch to be best-effort rather than guaranteed, since the requests run concurrently. For a nightly job with a stable system prompt, caching plus batching is the cheapest configuration available.

Is compaction the same thing as prompt caching?

No. Caching makes it cheaper to re-send context that has not changed. Compaction reduces how much context you send at all, by replacing older turns of a conversation with a server-written summary. They address different halves of the input bill, and a long-running agent will usually want both: a cached, stable prefix at the front and compaction keeping the growing history behind it in check.

Who should own cost control on a Claude API build: finance or engineering?

Both, with different jobs. Finance sets the ceiling, defines what "too expensive" looks like in pounds per month, and asks for per-request instrumentation as a delivery requirement. Engineers make the model routing, caching, batching, effort and compaction decisions that determine whether the system lands under or over that ceiling. The failure mode is one side owning the whole question: finance setting a budget with no visibility into what drives it, or engineers optimising for accuracy with no visibility into what it costs.

John Kelleher is a Claude Certified Architect (Foundations and Professional) and leads SpotDev, an OpenAI Select Partner that is also Certified in the Claude Partner Network.

Written by

John Kelleher

John is the founder and the Chief Executive at SpotDev.