Last checked 12 Sep 2026. Prices are converted from US dollars at £1 = $1.35 on 11 Sep 2026; vendors bill in dollars and UK VAT applies on top. HubSpot credit prices are HubSpot’s own UK list price in pounds.
Four separate clients raised the same point with us between July and September 2026, and none of them used the word “tokens”. One finance director put it as “our contact list inflates AI running costs”. A client-services team called the same problem “the noise”. They were right, and the arithmetic backs them up.
The model you pick sets a unit price. Your data sets the volume. On most CRM-connected agents the volume term is the larger of the two, and it is the one you control.
AI running cost scales with records touched and tokens per record
An agent that works inside a CRM has a simple cost shape: number of records it touches, multiplied by how many times it touches each one, multiplied by how many tokens it reads and writes each time. The model price is a multiplier on the end of that, not the driver of it.
Swapping a frontier model for a mid-tier one might halve your unit price. Cutting 4,000 tokens of raw account history down to a 400-token brief cuts the input side by ninety per cent, and it does so regardless of which vendor you are on. The two levers are not the same size.
The four cost multipliers in a messy CRM
- Duplicates. Every duplicate is a record the agent reads, reasons about and often contacts. A 15 per cent duplicate rate is a 15 per cent surcharge on everything the agent does, paid every month.
- Generic and personal mailboxes. info@, accounts@ and a director’s personal address all look like contacts. An outreach agent works them like contacts. You pay for the reasoning and the send, and nobody owns the reply.
- Missing owner and association data. When a contact is not linked to a company, a deal or an owner, the agent cannot take the short route. It re-reads the activity history to work out who this person is and what has already been said. That is the single most expensive failure mode we see, because it converts a lookup into a long read.
- Inconsistent phrasing. Five people describing the same product five ways means the prompt has to carry examples, exceptions and disambiguation. Longer instructions on every call, forever.
A worked cost model, with the assumptions stated
Assumptions, so you can substitute your own numbers: a 40,000-contact base with a 15 per cent duplicate rate, so 34,000 real people. An agent that reads a contact’s history and drafts a reply, touching each record four times a month. Before cleanup it reads 4,000 input tokens per record. After an overnight job writes a 400-token brief onto each record, it reads 400. Output is 300 tokens either way.
We have priced this on Anthropic’s Sonnet 5 at about £1.48 per million input tokens and £7.41 per million output, converted from the list price on Anthropic’s pricing docs, checked 12 Sep 2026. We use a mid-tier model here because that is where most production CRM agents land: cheap enough to run at volume, capable enough to be trusted on customer-facing drafts. The arithmetic is identical for any comparable mid-tier model, including GPT-5.6 Luna. Only the unit price changes.
| Monthly cost | Messy base | Cleansed and pre-summarised | With change-only refresh |
|---|---|---|---|
| Records touched | 40,000 | 34,000 | 34,000 |
| Agent reads (4 per record) | 160,000 | 136,000 | 136,000 |
| Input tokens per read | 4,000 | 400 | 400 |
| Agent input cost | £948 | £81 | £81 |
| Agent output cost | £356 | £302 | £302 |
| Overnight pre-summarisation | None | £302 | £30 |
| Total | £1,304 | £685 | £413 |
| Cost per record | 3.3p | 2.0p | 1.2p |
| Cost per conversation | 0.8p | 0.5p | 0.3p |
Two honest observations about that table. First, pre-summarising every record every month costs almost as much as it saves, because the summarising job has to read the same history the agent was reading. It only pays properly when you refresh the brief for records that actually changed, which in a typical base is around a tenth of them. That is the third column, and it is a 68 per cent cut.
Second, hygiene barely touches the output side. Output is the work itself: the agent still has to draft the reply. Anyone promising that a data cleanse will cut your whole AI bill tenfold is quoting the input line and hoping you will not ask. A tenfold reduction is achievable, but only where records carry very long histories and the agent re-reads them many times a day.
The same maths in HubSpot credits
If your agents run inside HubSpot rather than against a model API, the unit is credits, not tokens, and duplicates cost more visibly. HubSpot’s UK list price is £9.00 per 1,000 credits paying monthly and £8.10 paying annually. HubSpot’s published rate sheet charges the data agent 10 credits per prompt per record, and workflows 10 credits per AI action.
| One data-agent prompt across the base | Credits | Monthly rate (£9.00/1,000) | Annual rate (£8.10/1,000) |
|---|---|---|---|
| 40,000 contacts, 15% duplicates | 400,000 | £3,600 | £3,240 |
| 34,000 after deduplication | 340,000 | £3,060 | £2,754 |
| Cost of the duplicates | 60,000 | £540 | £486 |
An Enterprise subscription includes 5,000 credits a month, which covers 500 data-agent prompts. Everything above is overage. Set an account-level and a feature-level credit cap before you switch anything on, not after the first invoice.
One detail worth knowing before you budget: HubSpot’s own knowledge base (updated 13 Jul 2026) states that “Data enrichment is optional and is turned off by default in your HubSpot account.” So a base that has been sitting untouched for years is unlikely to have been filling its own firmographic gaps in the background.
The sequence that works
Run these in order. Reordering them is how AI budgets get spent twice.
- Cleanse and segment, as a priced workstream. Deduplicate, retire dead addresses, split generic mailboxes out of any outreach audience. Treat it as a project with a scope and a number, not as preparation somebody squeezes in. Merging at scale has its own traps, particularly on phone numbers: see deduplicating HubSpot contacts safely at any scale.
- Define owners and associations. Every contact linked to a company, a deal and a named owner. This is what stops the agent re-reading history to establish context, and it is usually the biggest single saving.
- Pre-summarise. An overnight job writes a short, structured brief to each record: who they are, what they bought, what was last said, what is outstanding. The agent reads the brief, not the archive. Refresh only what changed.
- Then switch the agent on. With a cap, a defined exit condition, and a per-record cost you can already predict.
How to measure it afterwards
Two numbers, tracked weekly from the day the agent goes live.
Cost per record is total monthly AI spend divided by distinct records touched. It should fall as hygiene improves and then flatten. If it rises, something is re-reading history it should not need.
Cost per conversation is spend divided by completed interactions. This is the one to put in front of a finance director, because it compares directly against the cost of a person doing the same work. It should be stable. A rising cost per conversation usually means the agent is looping, not that the model got dearer.
Log input and output tokens separately. Aggregate spend tells you the bill went up; the split tells you whether it was data volume or work volume. There is more on capping and monitoring model spend in our guide to controlling Claude API running costs.
When the answer is not to buy any AI at all
Sometimes the honest recommendation is that you do not have an AI problem.
If your contact base is 40 per cent duplicates, half your company records have no owner, and nobody trusts the reports, an agent will amplify all of that at a monthly cost. It will contact the same person three times, in three tones, on behalf of nobody. The cleanse is not a prerequisite to the AI project in that case. The cleanse is the project, and you may find that once it is done, the reporting and routing you actually wanted works without an agent in it.
We have told clients this and lost the AI scope as a result. It is still the right answer. The test is simple: if you cannot describe what good data looks like for the process you are automating, you are not ready to automate it.
Where to start
If the data is the problem, that is a data engineering engagement: deduplication, association repair, owner assignment and the pre-summarisation job itself. If you are further along and want the agent scoped against a running cost you can defend, start with the AI and Data Readiness Assessment on our AI implementation page, or take our diagnostics.
SpotDev is an AI and digital transformation consultancy that builds real software. We are a HubSpot Diamond Solutions Partner, Custom Integration Accredited, an OpenAI Select Partner and a Claude Registered Partner. To price a cleanse, a pre-summarisation job or an agent build, Request a Quote.
Frequently asked questions
Does choosing a cheaper model save more than cleaning the data?
Rarely. Moving from a frontier model to a mid-tier one typically halves the unit price. Cutting 4,000 tokens of raw history to a 400-token brief cuts the input side by about ninety per cent, and that saving applies whichever model you are on. Do both, in that order.
How much do duplicate contacts actually cost?
On HubSpot credits at the UK list price of £9.00 per 1,000, one data-agent prompt across a 40,000-contact base with 15 per cent duplicates costs £3,600, against £3,060 for the 34,000 real records. The duplicates cost £540 for that single pass.
What is pre-summarisation and why does it cut the bill?
An overnight job reads each record’s history once and writes a short structured brief to the record. The live agent then reads the brief instead of the archive. It works because a brief is read many times and written once. Refresh only the records that changed, or the summarising job costs about as much as it saves.
Which metrics should we track once the agent is live?
Cost per record and cost per conversation, weekly, with input and output tokens logged separately. Cost per conversation is the number to compare against the cost of a person doing the same work.
Does HubSpot fill in missing company data automatically?
Not unless you have turned it on. HubSpot’s knowledge base, updated 13 Jul 2026, states that data enrichment is optional and is turned off by default in your account.
Should we ever fix the data and skip the AI?
Yes, and it happens. If the records are bad enough that nobody trusts the reports, an agent will repeat those errors at a monthly cost. Do the cleanse, then re-test whether you still need the agent. Sometimes the routing and reporting you wanted works without one.
Stay Updated with Our Latest Insights
Get expert HubSpot tips and integration strategies delivered to your inbox.



