ChatGPT for operations leaders: when you need custom and when you do not

How to sort an operations process list into what a subscription covers, what a workspace agent covers, what needs building, and what AI should not touch.

John Kelleher
John Kelleher

You are not deciding about one process. You are looking at thirty or forty of them, somebody senior has asked where AI should go, and the answers coming back are a mixture of enthusiasm from people who have used ChatGPT at home and proposals from suppliers who have quoted before understanding the work.

The common failure is picking the most visible process, or letting whoever is selling do the picking. The useful move is to sort the list first. The sort is cheap, it takes an afternoon with the people who actually run the processes, and it is the one part of this that cannot be delegated to a supplier without a conflict of interest.

Five questions sort a process list faster than a workshop

Run every process on the list past the same five questions. You are not looking for a score. You are looking for which of the four buckets below it lands in.

QuestionWhy it decides anything
How often does it run, and how long does each run take?Something that happens twice a year will not repay a build, however painful it is each time. Frequency is what turns a saved half hour into a business case.
How structured is the input already?If a template works, you have a data-collection problem, not an AI one. AI earns its cost where the input resists templates permanently, because other people set the format.
What does a wrong answer cost, and who finds out?An internal draft that is wrong wastes ten minutes. A wrong answer to a customer, a supplier or a regulator does not. This sets how much checking you have to pay for.
Does anything downstream act on the output without a person reading it?This is the single largest cost driver on the list. A human reading the output is a free error handler. Removing them is a bigger change than it sounds.
Does it write into a system of record, and does that system have rules?Producing text is easy. Producing a value the CRM, ERP or finance system will accept, every time, is engineering.

Two things fall out of this immediately. Most processes fail question one or two, which means they belong to a person with a subscription. And a small number pass all five, which is where a build actually pays.

One number worth holding on to: on most sorted lists only a small minority reach the build tier. If a supplier's sort of your list produces eight, that is information about the supplier rather than about your operation.

The four buckets, in a line each

Tier one: a subscription and a person. One person doing one piece of work with the output in front of them. Nothing happens unless somebody makes it happen, and that person is silently doing all the error handling, which is exactly why removing them costs money. This covers more of a typical operations list than anyone selling AI services will volunteer, and if your honest sort puts most of the list here, that is the cheapest result available and usually the correct one.

Tier two: a workspace agent, configured by whoever owns the process. OpenAI introduced workspace agents on 22 Apr 2026, and availability is by plan, so check what yours includes. Somebody in your business creates and publishes one without an engineer: it runs on a schedule or on a trigger from another system, it keeps working when nobody is logged in, it works in Slack, it connects to common business apps, and it asks for approval before write actions by default. Runs are credit-metered rather than bundled into the licence, and on Enterprise an admin has to turn agents on, which is frequently why a business assumes this is out of reach.

Tier three: engineering. It begins where another system has to receive the output and act on it with nobody in the middle, where the output has to land in a system of record with enforced values, where the same input has to give the same output provably, where you have to prove afterwards what was decided on what input, or where the work reaches a system no connector covers.

Tier zero: the processes AI should not touch. Four kinds, and they are the next section.

The tier boundaries in detail, and what each one commits you to, are in what the OpenAI API does that ChatGPT cannot. Two things to do before you accept anybody's sort of your list. Publish one tier-two agent and run it for a month, because that month tells you precisely where it fails, and if it does not fail you have finished. And if you are a HubSpot business, run the list against the AI already in HubSpot as well, because prebuilt agents for well-trodden go-to-market work may already cover part of it with nothing to integrate.

Tier zero: the processes AI should not touch

Missing from most versions of this conversation, and the section that saves the most money.

Where checking costs as much as doing. If the only way to trust the output is for the expert to redo the work, you have added a step and called it automation. This catches more specialist processes than people expect.

Where the process is already broken, or the data underneath it is wrong. AI applied to a broken process performs the breakage faster and more consistently, and gives it a veneer of authority. If two of your systems disagree about who the customer is, that disagreement is the project. Fixing it usually delivers more than anything on this page.

Where the judgement is the product. Credit limits, pricing approvals, disciplinary decisions, anything where the point is that a named person took responsibility. Drafting support around those is fine. Automating the decision removes the only thing the process existed for.

Where it simply does not run often enough. A painful process that happens quarterly is a candidate for a better template, not a build.

What a sorted list actually looks like

Illustrative, not a recommendation for your business. The distribution is the point.

ProcessWhat decides itLands at
Drafting supplier and customer correspondencePerson present, varied inputTier one
Turning a long thread into a position before a meetingPerson present, output read immediatelyTier one
First draft of a tender or bid responseJudgement is the work, a person edits every lineTier one
Weekly operations report pulled from several placesScheduled, nobody present, humans read the resultTier two
Chasing an exception listScheduled, write actions with approvalTier two
Triaging a shared inbox into queuesFrequent, structured output, a person still acts on itTier two, unless routing must be enforced inside the ticketing system
First-line customer supportCustomer-facing, escalation design decides everything, see what is realistic for a UK contact centreDepends entirely on your documentation
Reading incoming orders and writing them into the ERPUnstructured in, enforced fields out, downstream actionTier three
Overnight reconciliation with exceptions raised automaticallyDownstream action, audit trail requiredTier three
Approving a credit limit or a price exceptionUnrecoverable, judgement is the productTier zero

Sequencing, and the portfolio problem nobody warns you about

Do the sort with the people who run the processes, not with a supplier in the room. Exhaust tiers one and two before commissioning anything, because both are reversible and both generate the specification you would otherwise pay someone to guess at. When you do commission, start with the process that has the clearest checkable right answer, not the one with the most visibility.

Then govern what you have created. Every agent needs a named owner, credit consumption needs to be visible to somebody who cares about it, and the whole set needs a quarterly review that retires what nobody uses. A hundred unowned agents is the unowned-spreadsheet problem at higher speed.

Both failure directions cost real money and they fail differently. Over-building leaves you with a project you did not need, which is the most expensive item available and at least visible on a budget line. Under-building leaves a process quietly depending on someone remembering to paste the right thing into the right window, which fails silently and usually at the worst moment.

The decision in front of you

Not which supplier, and not which platform. Sort your own list against the five questions, with the people who run the work, before you accept a quote on any part of it. You will find most of it is already covered by subscriptions you hold, a useful amount can be built by your own people this month, and one or two things are worth a proper engineering conversation.

If you have not yet decided which route you are on, how the OpenAI options compare for a UK business covers the ground.

We are an OpenAI Select Partner and a HubSpot Diamond Solutions Partner. We do not resell licences and we do not mark up your usage, so we have no reason to push a process up a tier, and our recommendation is more often to publish an agent and run it for a month than to build. If you want the sort done with you, including a straight answer on which items your existing subscriptions already cover, request a quote.

John Kelleher

John Kelleher

Author
John is the founder and the Chief Executive at SpotDev.

Stay Updated with Our Latest Insights

Get expert HubSpot tips and integration strategies delivered to your inbox.