ChatGPT for UK contact centres: what is realistic today

What AI genuinely does for a UK support team now, what is oversold, why your knowledge base is the real constraint, and why deflection is the wrong target.

John Kelleher
John Kelleher

You have probably been shown a number. Some proportion of your contact volume handled without a person, starting within weeks, with a payback calculation attached. The calculation is only as strong as the number, and nobody who produced it has read your documentation.

What follows is what a support operation can reasonably expect from this technology now, what the number usually hides, and the two design decisions, knowledge and escalation, that decide whether any of it survives contact with real customers.

Four jobs work today, and all four keep a person in the loop

Drafting replies. The model proposes, the agent edits and sends. Most of the value is speed on the routine phrasing that makes up the bulk of a queue, plus consistency from people who joined last month.

Summarising a history. A customer on their fourth contact about the same problem, with two channels and a returns note in the middle of it. The agent reads three minutes of history in fifteen seconds. This is the least impressive item on the list and usually the one your team rates highest.

Triage and routing. Classifying an incoming contact by type, urgency, product and tone, and putting it in the right queue at the right priority. Undersold, because a misrouted ticket costs a full handling cycle at both ends plus a customer who has to explain themselves twice.

Surfacing the right knowledge to a human. Pulling the passage that answers this specific question and showing it beside the conversation, with its source, so the person can check it before using it.

These four share a property. A person is the last checkpoint, so the worst case is a few wasted seconds rather than a wrong answer sent to a customer. That is why they are running in real operations now while full autonomy is still being argued about.

Autonomous resolution is real, but the band is narrower than the pitch

The honest version is not "AI cannot resolve contacts". It can. The band is questions with a single documented answer that does not depend on the customer's particular circumstances, plus a small set of actions the system has been explicitly configured to perform.

It stops at anything needing judgement, a commercial decision, an exception to policy, an apology that has to mean something, or an understanding of what the customer has not said. Those are not edge cases in a B2B support operation. On many desks they are most of the interesting volume.

The products themselves are built confidence-first rather than all-or-nothing, which is the correct design. HubSpot documents its customer agent this way: "If the customer agent knows the answer, it will respond to the visitor and provide relevant sources. If the agent doesn't know the answer, it'll ask the visitor to rephrase the question or transfer the conversation to a human agent."

The question to ask any supplier is not whether their agent can resolve contacts. It is where the confidence line sits, who is allowed to move it, and what happens on the wrong side of it.

Your knowledge base is the constraint, not the model

An agent answers from your content. HubSpot's customer agent syncs existing HubSpot content and crawls your site. So the ceiling on resolution is set by your documentation, not by the technology.

This is the part worth being blunt about. An agent answering from stale documentation is worse than no agent. It is confidently wrong at volume, with a citation attached that makes the wrong answer look checked, and the customer acts on it. A human reading the same out-of-date article hesitates, asks a colleague, or simply knows it is wrong because everyone on the floor knows.

So the first phase of the work is not a model decision. It is:

  • Last-reviewed dates on your top articles by contact volume, and a list of the ones nobody has touched since a product change.
  • The answers that exist only in an experienced person's head, or in a chat thread, and have never been written down.
  • Contradictions between your public site, your internal macros and what the team actually tells customers.

Stated plainly, an AI programme in a support function is a documentation programme with a model on the end of it. That work is a genuine cost, it is nearly always missing from the payback model you were shown, and if you cannot fund it you should not start. There is a compensation: every question the agent cannot answer becomes a ranked list of your documentation gaps, which is a better content roadmap than any audit will give you.

Deflection is the metric most likely to make things worse

Deflection counts conversations handled without human involvement. It sounds like success, and HubSpot's own documentation says why it is not: "Deflections do not always indicate resolution. For example, if the customer agent was unable to answer a question, the customer may have left the chat without requesting a human agent."

Read what that describes. A customer who gave up is counted as a win.

Now put a target on it. The levers available to a team asked to raise deflection are to make the exit harder to find, to delay the handoff, to ask one more clarifying question. Every one of those raises the number and degrades the experience. The cost does not disappear, it relocates: a repeat contact, a phone call, a complaint, a renewal conversation that goes differently, a review. Cost per contact falls while cost per customer rises, and only one of those is on the dashboard you built.

One signal worth noticing when you compare suppliers. HubSpot charges for this on resolution, not on volume handled: "Credits are consumed only when the agent delivers a resolution", where a resolution means a content-backed or action-based reply with no handoff inside 72 hours. Whatever platform you choose, ask what the vendor is paid for. A vendor paid on conversations contained has a different interest from yours.

Escalation design is what your customers will actually judge

Six rules, and none of them are technical:

  • The exit is available and obvious from the first turn. Hiding it buys a metric and loses a customer.
  • Context travels. The transcript, what the agent tried and what it could not answer all reach the human. The customer never starts again.
  • Handoff goes to a named queue with a service level, not into a general pool where it waits behind everything else.
  • The agent never promises what it cannot deliver. No refunds, no dates, no exceptions it has no authority to grant.
  • Escalation is rules-based as well as confidence-based. Named accounts, complaint language, anything with a safety or regulatory flavour, and anything where the customer has already asked twice, go straight to a person regardless of how confident the model is.
  • Unanswered conversations are reviewed weekly by someone who owns the content, not filed.

Defaults deserve a decision too. HubSpot's documentation says that if the visitor does not reply within 24 hours the chat closes automatically. It does not transfer to a person. That is the exact case HubSpot's own deflection caveat describes, arriving as a default rather than a choice, so decide what should happen to a silent chat rather than inherit it.

The internal half matters as much. If your team believes the agent is there to reduce headcount, they will not help improve it, and improving it is entirely dependent on them. If it visibly removes the first three minutes of every conversation, they will.

Measure resolution and satisfaction, then watch the queue you created

MeasureWhat it tells youThe trap
Resolution rateConversations fully resolved with no human handoffNot the same as deflection, and the gap between them is the whole story
Repeat contact within 7 days, on agent-resolved conversationsWhether the resolutions were realRises quietly while resolution rate looks good
Satisfaction, split three waysAgent only, escalated, human onlyA single blended score hides the escalated cohort, who are your most annoyed customers
Time to first human response for escalated contactsWhether you have improved the average by damaging the tailAlmost never watched, and almost always the thing that generates complaints
Volume of unanswered questionsNext quarter's resolution rateTreated as an error log rather than a backlog

If resolution rises while repeat contacts and escalated wait times also rise, you have moved cost, not removed it.

If you already run HubSpot, some of this is bought

For a HubSpot customer whose support sits in the conversations inbox, the honest answer is often that you do not need anyone to build you anything. HubSpot's customer agent is powered by Breeze, answers from your HubSpot content and public URLs, and is available on Professional and Enterprise tiers with credits consumed on resolution. HubSpot's own description of it is "Customer agent resolves common questions 24/7, with answers grounded in real contact and contract history, so your team is free for the conversations that need a human", and it now sits under Agent Hub, formerly Breeze Agents. That is a materially shorter route than a build, and what HubSpot's AI already covers is the first thing to check.

The same test applies on the OpenAI side. OpenAI's 22 Apr 2026 announcement put workspace agents on ChatGPT Business, Enterprise, Edu and Teachers, and you should check what your own plan includes today. They which run on a schedule or a trigger, work in Slack, and ask for approval before write actions, all configured by whoever owns the process rather than by an engineer. A fair amount of support-adjacent work sits inside that subscription: a weekly summary of contact themes, drafting responses for someone to approve, working an exception list.

Where engineering genuinely starts, in one line: you can trigger a workspace agent from another system, but OpenAI's documentation, checked 08 Aug 2026, states that "the agent's response cannot currently be retrieved through the API", so nothing can put the result back into your ticket automatically. Add the cases where a routing taxonomy has to be enforced rather than approximated, where the answer depends on systems no connector reaches (billing, dispatch, the platform your product runs on), where you need an audit trail on your own retention schedule, and where volume is large enough that how the calls are shaped changes the bill. Phone lines are a separate problem with their own constraints, covered in OpenAI's voice capabilities for UK phone and support lines.

The decision in front of you

Not "should we do AI in support". Two narrower ones.

Which of the four assisted jobs goes in front of your team this quarter, and who owns the documentation work that decides how far it gets. Then, before you switch anything customer-facing on, what you will be judged on, in writing, with deflection explicitly not on the list.

This is one rung of a longer ladder, set out in our overview of OpenAI for UK business.

We are an OpenAI Select Partner and a HubSpot Diamond Solutions Partner. Nothing we recommend earns us a margin on your platform spend, so we have no reason to talk you past the AI already inside the software you pay for, and in support that is frequently where this ends. If you want your knowledge base assessed for what it could realistically answer, and a straight answer on whether the platforms you already licence cover the rest, request a quote or read how we approach OpenAI implementation.

John Kelleher

John Kelleher

Author
John is the founder and the Chief Executive at SpotDev.

Stay Updated with Our Latest Insights

Get expert HubSpot tips and integration strategies delivered to your inbox.