AI Transformation7 Oct 202610 min read

Should Claude Haiku 5.5 or GPT-6 Luna run your volume work?

Claude Haiku 5.5 and GPT-6 Luna cost the same until 100,000 tokens. Where each wins on price, capability and UK residency, checked 7 Oct 2026.

Layered blue and orange tiers of light on a dark background

Checked 7 Oct 2026, the day Claude Haiku 5.5 launched. Prices are converted from US dollars at the rate used in our model decision table (1.35 dollars to the pound, set on 11 Sep 2026), and UK VAT applies on top.

The short answer: for short, demanding tasks, Claude Haiku 5.5 does more for the same price per token. For the cheapest possible volume, or anything that reads long documents, GPT-6 Luna is still the better buy. The two cost exactly the same until a prompt passes 100,000 tokens. After that, Haiku costs five times as much.

This year I have used GPT-6 Luna as my default model for building agents. It is cheap enough to run at volume, and I only move a step up to a larger model when a task needs it. Anthropic's release of Claude Haiku 5.5 on 7 Oct 2026 is the first time that default has had a real challenger at the same price. This post sets out where each one wins, using the vendors' own pricing and documentation and one independent benchmark.

What Anthropic released

Claude Haiku 5.5 is the fastest and lowest-priced model in Anthropic's current line-up, which also includes Claude Sonnet 5.5, Claude Opus 5.5 and Claude Fable 5.1. Anthropic positions it for high-volume, cost-sensitive work: summaries, classification, routing, document questions, live customer support, browser use and narrow tasks handed down by a larger model. It says Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding.

The practical specifications:

  • A 1M-token context window and up to 128,000 tokens of output, against 200,000 and 64,000 on Haiku 4.5.
  • Five effort settings, from low to max, with medium as the default. It is the first Haiku model you can tune this way, trading cost for depth of reasoning per task.
  • Computer use and browser use on the Claude API.
  • Availability on the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry.

Anthropic says it costs around 75% less to run than Haiku 4.5 on average. That figure blends a 90% cut on prompts up to 100,000 tokens with a 50% cut above, and allows for a new tokenizer that counts the same text as roughly 30% more tokens.

Same price per token, until 100,000 tokens

On short prompts the price lists are identical, including cached and batch rates. The difference is where each vendor puts its long-prompt surcharge.

Prompt sizeClaude Haiku 5.5 (per 1M tokens in / out)GPT-6 Luna (per 1M tokens in / out)
Up to 100,000 tokensAbout 7p / about 37pAbout 7p / about 37p
100,000 to 272,000 tokensAbout 37p / about £1.85About 7p / about 37p
Over 272,000 tokensAbout 37p / about £1.85About 15p / about 56p

Here is what that means for a thousand calls, each producing 1,000 tokens of output:

  • 20,000-token prompts: about £1.85 on either model.
  • 150,000-token prompts: about £57 on Haiku 5.5, about £11 on Luna.
  • 400,000-token prompts: about £150 on Haiku 5.5, about £60 on Luna.

Caching does not change the ranking. Both models read cached input at a tenth of the normal input rate, so caching cuts the bill on both by a similar proportion, and a cached 150,000-token prompt is still about five times dearer on Haiku.

Watch agent loops in particular. An agent's prompt grows with every tool result it reads, so a run that starts at 30,000 tokens can pass 100,000 by its tenth step. The price band is set by the largest prompt in the loop, not the first. If your agent reads long records, contracts or call transcripts, assume it will cross the line.

What the extra capability is worth

Anthropic's launch table compares Haiku 5.5 directly with GPT-6 Luna and shows Haiku ahead on every test where Luna has a score. Two of the larger gaps are in computer use (72.4% against 48.9% on OSWorld 2.1, offline subset) and agentic coding (39.2% against 16.4% on Terminal-Bench 4.0). These are vendor-reported figures on tests the vendor chose, so treat them as a reason to test rather than a result.

The independent picture points the same way, with a cost. On the Artificial Analysis Intelligence Index, both run at maximum effort, Haiku 5.5 scores 43 and Luna 38. To get there, Haiku used about 162,000 output tokens per task against Luna's 50,000. Output is the expensive half of the bill, so at that setting Haiku's higher score comes with a noticeably higher cost per task, even at identical rates. Lower effort settings will narrow the gap. Haiku defaults to medium, and Luna's own scores fall steadily as effort drops, so compare the two at the setting you will actually run.

Early customer reports are consistent with Anthropic's table. AlphaSense reported a statistically significant improvement over Haiku 4.5 on 400 document questions, and HubSpot said Haiku 5.5 got the best score it has seen on its simulated CRM test suite. Neither compared it with Luna.

Residency and platform differences

For a UK business this can decide the question before price does.

  • GPT-6 Luna can be processed in the EU directly from OpenAI, on Standard, Flex and Batch processing, at a 10% regional surcharge. OpenAI's UK region offers storage only, not processing.
  • Claude Haiku 5.5 on Anthropic's own API runs in the US or globally, with no EU or UK option. On Amazon Bedrock it has an EU geographic profile, which routes requests across European regions rather than keeping processing in the UK. At launch the AWS model card lists Bedrock's standard tier only, and Anthropic's newest computer-use toolset is not yet available there, so the EU route currently gives up some of the features above.

Neither model offers UK-only processing. If that is a hard requirement, see our guide to AI data residency in the UK.

Two operational differences are worth knowing before you build:

  • Discounted capacity. Luna has a Flex tier at batch prices for work that can wait, which may return a "resource unavailable" error at busy times, though OpenAI does not charge when it does. Anthropic offers batch processing at half price but no Flex equivalent. OpenAI also released a Decisions API in beta on 6 Oct 2026 for Luna, which it says turns text and images into typed answers ten times faster than its standard API. That is a vendor claim we have not yet tested.
  • Refusals. Haiku 5.5 runs safety classifiers that can decline a request, and Anthropic does not automatically retry a declined request on another model. Your software needs its own path for a refusal. Teams moving from Haiku 4.5 should also check their requests: assistant prefill, non-default temperature settings and manual thinking budgets now return errors.

How I would choose

My own rule, after a couple of hours with the documentation and before a full test programme:

  • Use Claude Haiku 5.5 when a typical request stays well under 100,000 tokens and the task is hard enough for capability to matter: multi-step tool use, computer use, strict instruction following, judgement calls on short documents. Start at medium effort.
  • Use GPT-6 Luna when price is the main driver, when requests regularly pass 100,000 tokens, when you want Flex pricing, or when you need EU processing direct from the vendor.
  • If both fit, route by prompt length, which you know before you make the call. Short and demanding goes to Haiku, long or routine goes to Luna.

Then test. Take twenty real examples from your own data, run them through both at the effort setting you plan to use, and compare the cost per finished task rather than the price per token. That takes about a day and tells you more than any table, including the one above. For the wider set of models and the criteria behind them, see our model decision table, and for the other levers on an AI bill, how to keep Claude API costs under control.

If you would like help choosing and building around the right model for your own work, our AI implementation team can run that evaluation with you, or you can Request a Quote.

Frequently asked questions

Is Claude Haiku 5.5 cheaper than GPT-6 Luna?

Not per token. Both cost about 7p per million input tokens and about 37p per million output tokens for prompts up to 100,000 tokens. Above that, Haiku 5.5 rises to about 37p and about £1.85, while Luna keeps its base price up to 272,000 input tokens, so Luna is cheaper for long prompts. Haiku also tends to generate more output tokens per task, so compare cost per finished task on your own work.

Is Claude Haiku 5.5 better than GPT-6 Luna?

On Anthropic's own benchmarks Haiku 5.5 scores higher on every test where Luna has a result, and on the Artificial Analysis Intelligence Index it scores 43 to Luna's 38, both at maximum effort. It used about three times as many output tokens per task to do so. For short, demanding tasks it is the stronger model. For simple, high-volume work the extra capability may buy you nothing.

Can Claude Haiku 5.5 or GPT-6 Luna keep data in the UK?

Neither offers UK-only processing. OpenAI processes Luna in the EU on request and offers UK storage only. Anthropic's own API runs Haiku 5.5 in the US or globally, and Amazon Bedrock offers an EU geographic profile that routes across European regions.

Should we move from Claude Haiku 4.5 to Haiku 5.5?

For most workloads, yes. Anthropic lists Haiku 4.5 for retirement not sooner than 15 Oct 2026, and Haiku 5.5 is a tenth of its per-token price on prompts up to 100,000 tokens. Test first, because Haiku 5.5 counts the same text as more tokens and rejects some request settings that Haiku 4.5 accepted.

Sources

Written by

John Kelleher

John is the founder and the Chief Executive at SpotDev.