Last checked: 12 Sep 2026. Covering model releases up to 8 Sep 2026. Every figure below comes from the vendor's own pricing or documentation page on that date, not from a comparison site.
The question we get asked most often is a version of this: "Do you use ChatGPT, or Claude, or something else, and why would you pick one over the other for us?" It is a fair question and it is usually answered badly, because most published comparisons rank models on coding tests and then present the ranking as procurement advice.
SpotDev is an AI and digital transformation consultancy that builds real software. We are an OpenAI Select Partner and a Claude Registered Partner, and we resell none of these models, so the table below is the same table we use internally. The working method is a ladder: put the cheapest model that can do the job on the volume work, and reserve an expensive model for the step where being wrong is costly. Most builds end up using two or three models, not one.
Prices are converted from US dollars at £1 = $1.35 on 11 Sep 2026. Vendors bill in dollars and UK VAT applies on top. HubSpot is the exception: it publishes a sterling price list.
The decision table, autumn 2026
| Option | Per 1M input / output tokens | Context window | Agent and computer use | UK and EU residency | Official HubSpot connection | Workload it fits |
|---|---|---|---|---|---|---|
| GPT-6 Astra (OpenAI, 3 Sep 2026) | About £7.40 in, about £37 out. Prompts over 272k input tokens bill at 2x input and 1.5x output. Fast mode 2x. Batch and Flex 50% off. | 1,050,000 tokens in, 128,000 out | Yes. Documented tools include computer use, MCP and a hosted shell, with effort settings from low to max. | EU region gives storage and processing. UK region gives storage only, so UK-resident inference is not available. Residency endpoints carry a 10% uplift. | ChatGPT connector. It creates and updates contacts, companies, deals, tickets, line items, products and engagements, up to 10 records per bulk action. | The hard step in a chain: final judgement, long-horizon agentic work, anything where a wrong answer is expensive. |
| GPT-5.6 Luna (OpenAI) | About 15p in, about 89p out. Cached input about 1.5p. | 1,050,000 tokens in, 128,000 out | Yes. Function calling, computer use and MCP. | Same residency list as Astra. EU storage and processing, UK storage only. | Same ChatGPT connector. | Volume. Classification, extraction, summarising, first-pass drafting. This is the workhorse rung of the ladder. |
| Claude Fable 5.1 (Anthropic, 1 Sep 2026) | About £7.40 in, about £37 out. Cache reads 2.5% of input price. Batch 50% off. | 1M tokens in, 128,000 out | Yes. Anthropic publishes per-tool token overheads for bash, code execution, computer use and browser use, plus a runtime charge for its managed agents. | No first-party EU or UK region. The API offers US-only routing at a 1.1x multiplier or global. EU hosting means routing through Amazon Bedrock or Google Vertex AI regional endpoints. | HubSpot connector for Claude, plus HubSpot's own remote MCP server. | Long-horizon reasoning and agentic work over large document sets. Tone-sensitive client-facing writing. |
| Claude Sonnet 5 (Anthropic) | About £1.48 in, about £7.40 out. | 1M tokens in, 128,000 out | Yes, the same tool surface as Fable 5.1. | As Fable 5.1. | As Fable 5.1. | The default middle rung. Most production steps that need judgement but not the flagship. |
| Gemini 3.8 Flash (Google, 2 Sep 2026) | About 56p in, about £2.78 out until 31 Dec 2026. From 1 Jan 2027 it doubles to about £1.11 and about £5.56. Batch 50% off. | 1,048,576 tokens in, 65,536 out | Google positions it for autonomous agents and long-horizon software engineering. It takes text, image, video, audio and PDF input and returns text only. | Google Cloud lists residency and processing in the US and EU multi-regions. It is global-only for in-country regions, which includes the UK (europe-west2). Gemini 3.5 Flash is the most recent Flash model supported in-country. | HubSpot connector for Gemini, in public beta. Read-only, and limited to contacts, companies, deals and their associations. | High-volume work where latency matters and a Google Workspace estate is already in place. Price it against the 2027 increase, not today's rate. |
| Grok 4.6 (SpaceXAI) | About £1.48 in, about £4.44 out under 200k context. About £2.96 and about £8.89 at or above it. | 500k tokens | Server-side Web Search and X Search tools, billed separately. X Search moves to per-post billing on 21 Sep 2026. | x.ai lists data residency under the Enterprise tier only, alongside no-training, custom retention, SSO, SCIM, audit and CMEK. The Business tier does not list it. | None published. | Work that genuinely needs live public-web or X data. Voice is a stated strength. |
| Grok 4.3 (SpaceXAI) | About 93p in, about £1.85 out under 200k context. About £1.85 and about £3.70 at or above it. | 1M tokens | Same search tooling, billed separately. | As Grok 4.6. | None published. | Cheaper long-context reading where the search tools are the point. |
| Muse Spark 1.3 (Meta, 2 Sep 2026) | About 93p in, about £3.15 out. A separate contributor variant is far cheaper at about 7p in and about 15p out, and its documentation states the data is used to improve Meta's products. | 1M tokens | The consumer Muse agent takes real-world actions from a secure VM, but it is US-only, 18+ and in limited testing. | None stated. | None published. | Hard to place in a UK B2B stack today. Watch it, do not build on it. |
| HubSpot Agent Hub (formerly Breeze Agents) | Priced per outcome in credits, not per token. £9.00 per 1,000 credits monthly, £8.10 annually. A resolved customer-agent conversation is 50 credits, so about 45p, or 40.5p on annual pricing. A prospecting-agent lead is 100 credits. A data-agent prompt is 10. | Not applicable | Agents act inside HubSpot. HubSpot's remote MCP server gives any MCP-compatible client read and write access to CRM data. | HubSpot hosts in regional data centres in the US, EU, Canada and Australia, and paid accounts can migrate to a preferred region. | It is the CRM. | Anything already fully inside HubSpot, where you would rather buy a finished outcome than build one. |
Two things the table cannot show. First, included credits: Starter accounts get 500 a month, Professional 3,000 and Enterprise 5,000, with no rollover. Second, the ChatGPT connector bypasses HubSpot validation rules on write, which is a governance decision, not a feature.
Eight criteria that actually decide it
1. Licences you already pay for
Before anything else, list what is already on the invoice. A Microsoft or Google estate, a HubSpot Professional or Enterprise subscription with monthly credits nobody is spending, an existing ChatGPT Business tenancy. The cheapest capable option is frequently one you are already funding, and saying so is often the honest answer to "which model should we buy".
2. Residency
This is the criterion most likely to remove options outright, and the position is messier than the marketing suggests. No first-party Claude inference runs in the EU or the UK. OpenAI gives the EU storage and processing but the UK storage only. Gemini 3.8 Flash covers EU multi-regions but not UK in-country. Grok lists residency at Enterprise tier only. If a regulator or an outsourcing policy requires in-country processing, most of this table fails that test today, and the buildable answers are a cloud provider's regional endpoints or keeping the data out of the model entirely.
3. Integration surface
What the model can reach matters more than what it scores. The gap between a read-only connector and one that writes back is the difference between a tool that answers questions about your pipeline and one that owns a stage of it. Right now the ChatGPT connector writes, the Claude connector writes, HubSpot's MCP server reads and writes, and the Gemini connector reads three object types. Grok and Muse have no published HubSpot route at all.
4. Agentic maturity
Judge this on documentation, not demos. Anthropic and OpenAI both publish the specifics: which tools exist, what computer use costs in extra tokens, how effort levels change behaviour. Where a vendor only markets the capability, treat the absence of documentation as an unknown rather than as a weakness, then test it before you commit.
5. Cost at scale, worked
Headline rates mislead until you run a real volume through them. Take 10,000 support conversations in a month, each averaging 8,000 input tokens of ticket history and knowledge-base context and 800 output tokens. That is 80M input and 8M output tokens.
- GPT-5.6 Luna: about £19
- Gemini 3.8 Flash: about £67 at the introductory rate, about £133 from 1 Jan 2027
- Grok 4.3: about £89
- Muse Spark 1.3: about £99
- Grok 4.6: about £154
- Claude Sonnet 5: about £178
- GPT-6 Astra or Claude Fable 5.1: about £889
- HubSpot customer agent: about £4,500, or about £4,050 on annual credit pricing
Read that honestly. The token figures are raw inference only. They exclude the retrieval layer, the platform, the engineering to build it and the human review step. The HubSpot figure buys a finished, supported product where you pay only on a resolved conversation. The comparison is not 45p against a fraction of a penny, it is "buy the outcome" against "build the thing". A forty-fold spread between the top and bottom rungs is the whole argument for a ladder, though. One client's per-record cost fell by roughly 90% when we had an overnight process pre-summarise account history instead of making the agent re-read everything live, with no model change at all.
6. Vendor churn risk
Anthropic's own pricing page lists seven legacy Claude models still available alongside four current ones. OpenAI's GPT-5.6 family was superseded within weeks. Google shipped three Flash releases in six weeks, and its own model card for 3.8 Flash says the safety evaluation found no meaningful new capabilities against 3.7 Flash. The OpenAI Assistants API was sunset on 26 Aug 2026 and is no longer available. This is not a reason to pick one vendor over another. It is a reason to keep the model name in configuration, behind an interface, so swapping it is an afternoon and not a project.
7. The shape of the pricing model
Three shapes are on offer. Per-token pricing suits variable and unpredictable volume. Per-seat pricing suits universal adoption where everyone uses it daily. Per-outcome pricing, which HubSpot is alone in offering here, suits low-volume high-value tasks and makes budgeting trivial. The shape usually matters more than the rate, because it decides whether your bill scales with usage, with headcount or with results.
8. Speed
Latency is a product decision, not a benchmark. A customer waiting on a live chat reply and a nightly batch job have opposite requirements. Anthropic ranks Sonnet 5 fast and Fable 5.1 slower within its own lineup. Google positions Flash as the speed tier. Both OpenAI and Anthropic charge more for faster or higher-effort modes. Decide the acceptable wait first, then shop within it.
What the leaderboards are actually measuring
Public benchmark leaderboards are coding and reasoning tests repackaged as business comparisons. They are real tests, honestly run, and they were never designed to answer "which of these should handle our quote chasing". Nothing on them measures whether a model can find the right record in your CRM, respect a permission boundary, or hand over to a person at the right moment.
The gap between vendor-reported and independent scores is the clearest illustration. OpenAI reported 99.9% for GPT-6 Astra on ARC-AGI-3. ARC Prize's own published result shows 62.7% under its standard harness and 99.9% only through OpenAI's provider adapter, which preserves opaque reasoning state between requests and compacts long conversations so the model can reuse prior work. Both numbers are true. They measure different things, and only one of them is the model.
Then look at the ranking itself. On the Artificial Analysis Intelligence Index, checked 12 Sep 2026, Claude Fable 5.1 and GPT-6 Astra both score 53. Gemini 3.8 Flash and Grok 4.6 are not in that index at all. A tie at the top and two absentees is not a procurement shortlist. Use benchmarks to exclude models that are obviously unfit, then run your own evaluation on twenty real examples from your own data. That evaluation takes about a day and it beats every leaderboard for your specific decision.
How this page is maintained
This table is updated within three working days of each major model release. Prices, context windows, residency positions and connector capabilities are re-checked against the vendor's own page at each update, and the "last checked" date at the top moves only when that has actually been done.
- 12 Sep 2026: first published, covering releases to 8 Sep 2026.
How we would use this with you
Model choice is the last decision, not the first. In most engagements the binding constraint is data quality, a permissions model or an integration that does not exist yet, and the model is a line item. Our approach across AI implementation is to settle the workload, the residency position and the approval step first, then pick the cheapest rung that clears the bar. Where a single vendor is the right answer we go deep on it, through OpenAI implementation, Claude implementation, SpaceXAI implementation or HubSpot AI. For a two-vendor comparison in more depth, see Claude vs ChatGPT for business.
If you want a view on your own situation rather than a table, start with a diagnostic, or Request a Quote.
Frequently asked questions
Which AI model is cheapest for a UK business in autumn 2026?
On raw API pricing, GPT-5.6 Luna at about 15p per million input tokens and about 89p per million output tokens is the cheapest option in this table. Gemini 3.8 Flash is next at about 56p and about £2.78, though that is an introductory rate that doubles on 1 Jan 2027. Prices are converted from US dollars at £1 = $1.35 on 11 Sep 2026 and UK VAT applies on top.
Can I keep AI processing of our data inside the UK?
Not with most first-party model APIs today. OpenAI's UK region offers storage only, with processing in the US or the EU. Google lists Gemini 3.8 Flash for EU multi-regions but not for UK in-country regions. Anthropic offers no first-party EU or UK region, only US-only or global routing, with EU hosting available through Amazon Bedrock or Google Vertex AI. If in-country processing is a hard requirement, design for it explicitly rather than assuming a vendor tier covers it.
Which AI assistants can write back to HubSpot?
As of 12 Sep 2026, the ChatGPT connector and the HubSpot connector for Claude can both create and update CRM records, and HubSpot's remote MCP server gives any MCP-compatible client read and write access. The HubSpot connector for Gemini is read-only and covers contacts, companies and deals. There is no published HubSpot connector for Grok or Muse.
Should we standardise on one AI vendor?
Fewer vendors means fewer data-residency questions to answer and fewer contracts to govern, which is a real benefit. It is also usually more expensive, because you end up running volume work on a flagship model. A reasonable middle position is one approved vendor for anything touching customer data, with a cheaper second model allowed for internal, low-risk, high-volume steps.
How much does HubSpot's Agent Hub cost compared with building an agent?
HubSpot charges per completed outcome in credits, at £9.00 per 1,000 credits monthly or £8.10 annually. A resolved customer-agent conversation costs 50 credits, about 45p. A custom build's inference cost for the same conversation is a fraction of a penny, but that excludes the platform, the retrieval layer and the engineering. Agent Hub is usually better value at low volume or where the work sits entirely inside HubSpot. A build wins at high volume or where the process crosses systems.
Do benchmark scores tell us which model to buy?
No. Public leaderboards are coding and reasoning tests, and vendor-reported scores can differ sharply from independent ones. OpenAI reported 99.9% for GPT-6 Astra on ARC-AGI-3, while ARC Prize's own standard harness gives 62.7%. Use benchmarks to rule out obviously unfit models, then evaluate the shortlist on twenty real examples from your own data.
Stay Updated with Our Latest Insights
Get expert HubSpot tips and integration strategies delivered to your inbox.




