Before an AI answers a single customer, somebody has to write down which questions it may answer and which it may not. One page, named categories, no principles. Where a pilot has stalled, that page is usually missing. The demo went well, then somebody asked what happens if it promises a refund, and the project has been waiting for an answer ever since.
That page is a policy decision, not a technical one. It belongs to whoever is accountable for the customer relationship rather than to whoever configures the tool, and it has to exist before anything goes live.
The boundary runs along two axes, not one
Most teams draw the line by topic alone. Billing questions yes, contract questions no. Topic is half the decision. The other half is what the reply does, because "here is where the returns policy sits" and "your return has been approved" are the same topic and completely different exposures.
So record two things per workflow: the topic, then the action, taken from a fixed ladder that runs from read at one end to restricted at the other. Permission attaches to the action, not to the agent. One capability can sit at bounded autonomous for sending an invoice copy the customer is already entitled to see, at act with approval for moving a delivery date, and at restricted for anything that changes what the customer pays. That is one system carrying three permissions, which is the only version that survives a real service desk. Higher positions are earned from evidence in live operation, and none of them removes the named person who owned the workflow.
What is escalated from day one and stays escalated
Some categories go to a named person from the start, written as restricted before launch and not relaxed later.
- Refunds, credits and goodwill. Any amount. A £12 credit sets the same precedent as a £1,200 one, and the customer will quote it back to you.
- Complaints, including the ones nobody labelled as complaints. They arrive as the fourth message in a thread, with the tone changing.
- Anything legal or contractual. Retrieving a clause and showing it is one thing. Telling a customer what that clause means for their situation is not a service task.
- Account closure, cancellation and downgrade. A customer asking to leave is telling you something the business needs to hear from a person.
- Any sign of vulnerability or distress. No automated reply, in any tone, at any confidence level. It routes to a person immediately.
- Requests about a customer's own personal data. These go to a named owner, and what your business owes in response is a question for your own advisers.
Write those as categories with worked examples, not keywords: a keyword rule catches "cancel my order" and misses "we have decided not to continue". Classify by what the customer is trying to achieve, and treat an uncertain classification as a routing instruction rather than a prompt to improve.
Deciding that something must reach a person is only half the job. What travels with it is what the customer actually experiences, and that is an engineering question we have written up separately in what has to move when an AI hands a customer to a person.
Permission to draft and permission to send are two separate grants
The most common design fault is treating "it can compose the reply" and "it can send the reply" as one switch. They are different decisions, and granting them together by default is how an organisation ends up unable to say who authorised a message that went out at 2am.
A drafting grant is cheap to withdraw, because every output passes a person first. A sending grant is not, and it should specify the channel, the audience, the message types, the operating hours, a rate limit and whose name appears at the bottom. The test before granting it: when this message is wrong, who is on the hook, and did that person agree to be? Two defaults hold permanently: read access grants nothing outbound, and an agent does not open a conversation with a customer who has not opened one first.
A confidence score is not evidence
This is the point most teams get wrong. A model's confidence score describes the model's own state. It is not a measurement of whether the answer is correct, and certainly not of whether the answer applies to the customer asking.
A customer asks whether their contract covers out-of-hours support. There is a published article saying it does. The retrieval is clean, the source is current, the wording is unambiguous, and any sensible system reports high confidence. The answer is still wrong, because that customer is on the older tier where it is not included. Nothing in the confidence figure could have caught that, because nothing about it was uncertain.
Gate on two things instead. Source: can this answer be traced to a specific approved source, and does that source apply to this customer's contract, region and entitlement? A general article is a valid source for a general question and an invalid one for an account-specific question. If the answer depends on which customer is asking and that cannot be established, it does not answer. Action type: what does the reply do? An answer pointing at published information sits in a different category from one stating a commitment, and the category is fixed in advance by policy rather than judged per message.
Confidence works as a secondary filter behind those gates. On its own it fails fluently and without hesitation. The related trap is containment rate, the share of conversations closed without a person, which rewards a system for guessing. Tune for the cost of being wrong instead.
Where HubSpot's own agent is enough
If you run HubSpot, part of what you are about to buy may already sit in your subscription. Sometimes the answer is the AI you already pay for, and we would rather say so than sell round it.
HubSpot's customer agent responds to customer questions "using your existing content", and also "performs configured actions (such as resetting a password or booking a meeting)". When it cannot answer, HubSpot documents that "it'll ask the visitor to rephrase the question or transfer the conversation to a human agent", and that it transfers when "its confidence is low or escalation rules are met". Both quotes are from HubSpot's knowledge base, retrieved 16 Aug 2026.
That is enough when the questions have settled answers living in published content, the knowledge base is current, and a wrong answer produces a mildly irritated customer rather than a credit note. If that describes your desk, configure it properly and spend the engineering budget elsewhere.
Two boundaries decide whether it stops there. It answers from content, so it does not know an order status sitting in your ERP or why an invoice was raised twice, and reaching those is an integration problem. On timing, HubSpot states that knowledge base articles re-sync when edited while "other content sources are automatically re-synced on a weekly basis", which is fine for a product overview and not for lead times or prices. What your own subscription includes is account-specific and changes, so check yours.
Draft for a human is a destination, not a stage
Drafting for review is not a training-wheels phase you graduate from once trust is established. For a large share of customer-facing work it is the correct permanent setting, and moving off it would make the workflow worse.
It is the right end state wherever one wrong message is expensive, wherever the answer depends on commercial judgement rather than fact retrieval, and wherever the volume is low enough that human review costs almost nothing anyway.
The value in those workflows is real and it is not autonomy. The reply arrives assembled and grounded in seconds instead of after twenty minutes of hunting through four systems, and a person who knows the account reads it and sends it. Time to first draft falls, judgement stays where it was, and that is the intended shape of AI customer service built inside HubSpot.
How you would know the boundary is set correctly
Four measures, and none of them is containment.
- Restricted-category messages that reached a customer without a person. The target is zero and it is checkable in the log.
- Correction rate on drafts. Repeated edits in one category mean the source is wrong or the category is misclassified.
- Escalation precision, both ways. Sample escalations that did not need a person, and handled conversations that did.
- Time to first draft. The measure that captures the benefit in a draft-permanent workflow.
Baseline all four before anything is built.
The next step
Writing the boundary needs whoever owns the customer relationship and whoever will be accountable for the workflow in the room. It is the first step of the customer success wave in our 12-month AI programme, which brings one department onto AI per quarter, before a line of configuration.
If you want an outside read first, start with a diagnostic. The AI and Data Readiness Assessment is short, fixed in scope and carries no obligation. For service work it will tell you which workflows can carry an agent and which should stay with a person. The question above this one is sequencing, which we cover in how to roll out AI department by department.
Stay Updated with Our Latest Insights
Get expert HubSpot tips and integration strategies delivered to your inbox.




