Which Admin Work AI Should Handle, and Which It Should Not

A risk framework for administrative work: which tasks carry legal, financial or reputational consequence and keep a named human owner, and which do not.

John Kelleher
John Kelleher

The standard advice is to automate the boring bits. It is the wrong test, and following it puts your largest exposure in exactly the place you are not looking.

Repetitiveness measures how often a task happens. It says nothing about what it costs when it goes wrong, who finds out, or how long the damage takes to surface. Only the second set of questions should decide what you hand to software. What follows is how we classify administrative work before building anything, and it is worth doing whether or not you build.

The most repetitive work is often the most consequential

Consider why a task became repetitive. Work turns into a routine because somebody wrote a procedure for it, and procedures get written for things that matter. Nobody documents a step-by-step method for work nobody would notice going wrong. So repetition is frequently a signal of consequence rather than the absence of it.

The pattern is plain across a mid-market back office. Payment runs are highly repetitive and irreversible. Supplier bank detail changes are repetitive, and a wrong one sends your money to somebody else. Meanwhile some of the least repetitive work in the building, the one-off internal summary nobody will read twice, carries almost no consequence at all.

Sort your admin by tedium and you reach for the first list. Sort it by what happens when it is wrong and you get a different answer.

Classify by consequence, not by tedium

Four bands, and every administrative workflow lands in one of them. The band sets the default position, and the default holds until evidence moves it.

BandWhat it looks likeDefault position
Contained A wrong output is visible immediately and costs nothing outside the team. Collating, formatting, summarising for internal reading Delegate, with monitoring
Recoverable A wrong output causes delay or rework and can be put right in the ordinary course of business. Classification, routing, task creation, internal chasing Delegate, with an explicit failure route and a queue you can see
Externally visible A customer, supplier, bank, auditor or regulator sees it. A wrong output damages a relationship, and you may never be told Drafted by the system, sent by a named person
Irreversible or controlled Money leaves. A controlled record changes. A commitment is made, terms change, personal data is disclosed, or the tone of a customer conversation escalates Restricted. Do not automate. A named person authorises, and the system prepares the evidence

Two things about that fourth band, because it is where governance conversations usually go vague. Nothing an agent does may determine VAT or any other tax treatment, post to a ledger or other controlled financial record, make or schedule a payment, issue a write-off or credit, or agree a payment arrangement. That is not caution for its own sake. An agent that cannot post to your ledger cannot post the wrong thing to your ledger, and the constraint is what makes the rest of the workflow safe to run in production.

And to be plain: we give no financial, tax, legal or accounting advice. Where a classification question belongs to your accountant, auditor or solicitor, that is what we will tell you, and what we build reflects the boundary they set rather than one we invented.

Four questions that place a task in about a minute

Who sees the output if it is wrong. An internal colleague will tell you. A customer usually will not, and a supplier certainly will not if the error is in their favour.

How long before anyone notices. The dangerous combination is low visibility and slow detection. A misrouted internal task surfaces in a day. A wrongly categorised supplier record surfaces at year end, if you are lucky.

What putting it right costs, and whether it can be put right at all. Some things are a correction. Some things are a payment you are now asking a supplier to return as a favour.

Whose name is on it when someone asks. People skip this one, and it is the most useful of the four. If nobody can answer it, you have found a finding, and you found it before any software was involved.

Accountability does not follow interest

The temptation is to hand over dull work wholesale, precisely because nobody wants to own it. That absence of an owner is not a reason to automate. It is usually the reason the work is already a risk.

Human accountability for a workflow is never removed. What changes is where the accountable person spends their attention: on exceptions and on approving the consequential steps, rather than on performing the mechanical ones. A named person, not a team and not a distribution list, because a workflow owned by everybody is owned by nobody at the moment it matters.

The mechanism that makes this workable is granting autonomy per action rather than per agent. Every action is placed on a written scale before release, and higher autonomy is earned through evaluation and operating evidence rather than assumed at launch.

That distinction does more work than it looks. One workflow can legitimately run at three levels at once: read the shared mailbox freely, draft the reply automatically, and stop dead before anything is sent outside the company. Treating autonomy as a property of the agent forces one setting for the whole thing, which is always either too loose at the top or useless at the bottom.

Some admin should be deleted, not automated

The report nobody opens. The field maintained because a colleague who left in 2023 asked for it.

Automating those is worse than leaving them alone. Automation makes them cheap, and cheap things stop being reviewed. You have converted a visible annoyance that somebody might eventually kill into a permanent invisible one.

The test is unglamorous and it works: stop it for a month and see who asks. If nobody asks, you have your answer.

That test applies to administrative steps only. Anything forming part of a financial control, an audit trail or a regulatory record is out of scope for it by definition, and the person who owns that control decides whether it can change, not the person who finds it annoying. A test that tells you to switch off a control is obviously wrong, which is why this one names its own limit.

What we would actually build

Classification and routing in a shared mailbox, with the account context and a draft reply attached, and external sending held behind a separate authority. Internal request handling that gathers what a request is missing, creates the tasks and tracks the non-responses. Queue monitoring where anything unresolved past a threshold you set appears with its age against it. Document intake, where the extraction is checked and approved by a person rather than posted.

What distinguishes this from a bolt-on is not the capability list, which any vendor can match. It is that the work happens where your records already live. That architecture brings its own copy of your data, its own login and its own definition of what counts as a customer, so two systems can disagree about who you are dealing with while both look authoritative. Built into the systems you already run, the request, the record, the approval and the audit trail are one thing, and a question about who approved what has one answer instead of a reconciliation. These sit in our AI for Business Operations pack, and a standard wave activates two of them rather than the whole catalogue.

How you would know it worked

Hours saved is the wrong measure, partly because nobody can verify it and mostly because saved time is not automatically cost reduction. Better measures, all from your own systems rather than a survey:

  • How many items sit unresolved past their threshold, and how long the oldest has been there.
  • How many touches an item takes from arrival to closure.
  • How often work reaches the wrong owner and has to be rerouted.
  • How many externally visible drafts are corrected at the approval gate, which tells you whether the gate is earning its place.
  • How often a workflow has to be moved into a stricter band after release, which tells you whether your classification was honest.

Capture those before anything is built. A baseline taken afterwards is a negotiation, not a measurement. That sequencing is why the AI Accelerator programme runs a quarter at a time, department by department, with every wave ending in a decision that includes stopping.

The next step

Run the four questions across your own workflows before looking at any tooling. Look in particular for work in the fourth band that somebody has already half-automated, and work in the first band that several people are still doing by hand.

If you would rather do that with us, it is a short, fixed-scope assessment with no obligation. Book a diagnostic and we will map the work by consequence, tell you what is genuinely safe to delegate, and say which parts we think should simply stop. The same classification drives how to order a wider AI rollout.

John Kelleher

John Kelleher

Author
John is the founder and the Chief Executive at SpotDev.

Stay Updated with Our Latest Insights

Get expert HubSpot tips and integration strategies delivered to your inbox.