John Kelleher is a Claude Certified Architect (Foundations and Professional) and leads SpotDev, a Claude Registered Partner and OpenAI Select Partner.
You have been asked to "do something with AI" and you have a budget line, a vendor demo and no way of knowing, in three months, whether it worked. The fix is not a bigger programme. It is one workflow, one pilot, a written definition of success before anyone logs in, and a decision at the end you can defend to the board.
The three-step shape behind this post (foundation, pilot, scale) and the pilot timeline come from Anthropic's Enterprise AI Transformation Guide (Oct 2025). The guide is written for enterprises with steering committees and analytics teams. What follows is that framework cut down to a firm of 30 to 500 staff, tested against independent evidence, with the gaps named. The 8 to 12 week window is our recommendation, informed by the guide. You will see the same window attributed to analyst firms; we could not find a primary source for that and do not repeat it.
What a pilot has to prove
A pilot has one job: to produce enough evidence, on your own work and data, to make one of three decisions.
- Scale it. The gain is real, measured against a baseline, and worth the cost of extending it to more people or more of the process.
- Keep it running at this size. It works for the five people using it and there is no case for spending more. This is a legitimate outcome, and in our experience it is the most common one in a firm of your size. None of the enterprise frameworks name it; their whole frame is pilot-to-scale.
- Stop it. It missed the criteria you wrote in week 1. The money spent bought a clear answer, which is what a pilot is for.
Set expectations with UK evidence rather than vendor decks. The British Chambers of Commerce found 54% of UK SMEs actively using AI in 2026, up from 23% in 2023. The ONS found that only 10% of adopting businesses describe their use as extensive, and the average adopter uses between 1.4 and 1.6 AI technologies. DSIT's 3,500-interview study found 77% of adopting businesses reported no change in revenue and 12% reported an increase. Adoption is wide and shallow. A pilot that ends with a measured result puts you in a small group.
Choosing the one workflow
Pick one workflow. Not a department, not a portfolio. Four tests, all of which must pass.
Inside the frontier. The task must be one where current models are reliably good: drafting from a template, summarising a defined document set, classifying and routing, extracting fields, first-line answers from an approved knowledge base. In a study of 758 consultants, those using AI on tasks inside the model's competence completed 12.2% more tasks, 25.1% faster, with over 40% higher quality. On a task outside that competence, AI users were 19 percentage points less likely to get the right answer. A pilot on the wrong side of that line fails for reasons that have nothing to do with your team.
Cheap to be wrong. The guide advises against customer-facing or mission-critical work for a first pilot. We agree, with one refinement: the test is the cost of a wrong output, not whether a customer is involved. A draft reply a human reads before sending is cheap to be wrong. A price quoted with no check is not.
An owner who carries the risk. One named person who runs the workflow today, will use the tool daily, and is judged on its output. Not the IT lead, not the MD. If nobody will own it, the pilot has already failed the adoption test.
Data that can be matched. You must be able to say, per item of work, how long it took and whether it was right, before and during the pilot. If the work is in HubSpot (tickets, deals, sequences, notes) that is usually possible. If it lives in inboxes and heads, fix that first. Our post on AI readiness assessment covers what to check before you start, and the AI transformation ladder covers which type of tool to pilot first.
Where to look: gains are largest where work is high volume, repetitive and done by less experienced staff. In a study of 5,179 customer support agents, an AI assistant raised issues resolved per hour by 14% on average and 34% for novice and low-skilled agents, with minimal effect on experienced ones. The UK Department for Business and Trade's evaluation of 1,000 Microsoft 365 Copilot licences found time saved on drafting, summarising research, transcribing meetings, searching for information and brainstorming, and none on image generation, scheduling or producing slide decks. Pilot the first list, not the second.
Weeks 1 to 2: set up the test
Nothing is built in the first fortnight. Five things happen, in order.
Write the go/no-go criteria. One page, signed by the owner and the sponsor. For each of the four dimensions below, a number that means "scale", one that means "keep as is" and one that means "stop". Write the stop numbers first. It is easier to agree what failure looks like before anyone is attached to the tool.
Run a baseline week, measured by hand. The pilot users do the workflow exactly as now and for one week record three things per item: minutes taken, whether it needed rework, and volume. A shared spreadsheet is enough. A baseline captured after the build is a memory, not a measurement. Our post on measuring whether AI is working makes the general case; this post applies it to a pilot.
Name a hold-out group. Two or three people who do the same work and do not get the tool until the pilot ends. They keep recording the same three numbers. Without them you cannot separate the tool's effect from a busy month, a quiet month or the novelty of being watched. The academic studies use control groups and blind grading; the vendor guides use dashboards. A hold-out group is the cheapest version of the rigorous method a firm your size can afford.
Complete a DPIA before real customer data goes anywhere near the tool. If the pilot processes personal data in a new way (a tool reading customer records, transcribing calls, drafting replies from a CRM), UK GDPR requires a data protection impact assessment before the processing starts where new technology is likely to create high risk, and the ICO names AI as an example of that technology, so the assessment belongs before the first real record is read, not once the results look promising. The DPIA is also where you settle the vendor's data terms: retention, training use, sub-processors and the region data sits in.
Sort access and permissions. The tool gets the narrowest access that lets the workflow run. Read-only wherever possible. A named service account, not a person's login. A log of what it read and wrote. If it will write to HubSpot, decide now which properties it may change and put a human approval step in front of anything a customer will see.
Weeks 3 to 8: build, first users, weekly check
Build against the workflow as it actually runs, not as the process document describes it. The first real users go on in week 3 or 4, owner first. Expect the first fortnight of use to be noisy; the guide's own timeline allows two to three weeks for onboarding before efficiency effects show.
Every week the owner fills in one row against the four dimensions, using the smallest measurable proxy a spreadsheet plus the tool's own usage export can produce. No new tooling.
| Dimension | Smallest measurable proxy | Where it comes from |
|---|---|---|
| Adoption | Number of pilot users who used the tool on three or more days this week | The tool's usage export (every serious product has one) |
| Efficiency | Median minutes per item on a sample of ten items per user, against the baseline week and the hold-out group | The same spreadsheet used in the baseline week |
| Quality | Rework rate on a random sample of ten outputs, graded by someone outside the pilot team | A ten-minute review slot, weekly |
| Satisfaction | One question to each user: "Would you object if we removed this on Monday?" Yes or no. | A message, not a survey tool |
Three notes. Adoption from the tool's own export beats asking people. The efficiency number only means something next to the hold-out group's number for the same week. Quality is graded by someone who did not build or champion the tool: the NIST AI Risk Management Framework calls for assessment by internal experts who were not front-line developers, or by independent assessors. In a 100-person firm that means the finance manager or a team lead from another department spends ten minutes a week grading a sample. It is the cheapest control against a pilot team marking its own homework.
If a stop number is hit in this window, stop. Do not extend to see if it improves. The guide is right that when results are absent you "adjust your approach or use case, not your timeline".
For the mechanics of building a test set that catches the failures that matter, see how to test an AI agent. This post does not repeat it.
Weeks 9 to 12: volume, independent check, decision
The volume test. For at least two weeks the tool handles the full volume of the workflow for the pilot users, not a curated subset. Most pilots that look good in week 6 were being fed the easy items. Full volume is where the edge cases arrive: the malformed document, the customer writing in a second language, the ticket that belongs to a different process. Count them. They are the cost of scaling, and you want that number before you decide.
The independent spot-check. In week 10 or 11, someone outside the pilot team grades a random sample of thirty to fifty outputs from the volume weeks against the same quality standard used all along. The pilot team does not pick the sample. If the spot-check disagrees with the weekly numbers, the spot-check wins, and the disagreement is itself a finding.
The decision. In week 12 the owner and sponsor take the week 1 criteria and the twelve weekly rows and pick one of the three outcomes. Write the decision down with the numbers that drove it. If the answer is "scale", the next question is which department goes next, and our post on rolling out AI department by department covers that. What happens after go-live, including the failure modes that only appear in production, belongs to why AI pilots fail in production and monitoring AI systems. This post stops at the decision.
What moves the cost
We do not publish a pilot price in a blog post. The range is wide and the number depends on five things you can assess before asking anyone for a quote. For context, the Lloyds Bank Business Barometer found the largest group of UK AI investors (33%) had spent under £25,000 in total, with 18% spending between £25,000 and £100,000. That is total AI spend, not a pilot figure, and it says most firms are running modest experiments rather than programmes.
Data access. Is the workflow's data in one system with an API, such as HubSpot, or in five places including a shared drive and an inbox? Every extra source adds connection work and a data-quality problem.
Write permissions. A tool that reads and drafts is cheaper and safer than one that changes records or sends messages. Each write action needs a rule, an approval step and a rollback.
Integration depth. A pilot can start with a person copying between the tool and the CRM. Slower per item, far cheaper to build, and often the right first step. Deep integration is a scaling cost, not a pilot cost.
Review load. Every output a human must check costs minutes. If the quality bar means checking everything, the tool can work and still fail the efficiency test. Decide the review rule up front.
Seats and usage. Most tools price per seat, per outcome or per unit of usage. Five users for twelve weeks is a different commitment from fifty. Model the pilot cost and the scaled cost separately, and do not sign for the scaled number to get the pilot.
If you want a scoped figure for your workflow, the diagnostic is where we produce one.
How to stop a pilot cleanly
Nobody writes about this, and it is the part a finance director needs most.
Kill criteria are written in week 1. Three or four conditions, any one of which ends the pilot. Typical shapes: adoption below the "stop" number for two consecutive weeks after week 6; the independent quality check below standard; no efficiency gain against the hold-out group by week 8; a data incident of any kind. Writing them early is what makes stopping a decision rather than a failure.
Contain licence and contract exposure before you start. Monthly terms, not annual. Seat counts sized to the pilot group, not the department. A written route to export and delete your data on exit. If the vendor will only sell an annual commitment for a pilot, that is a finding about the vendor.
Tell staff in a way that does not sour the next attempt. Say what the pilot was testing, which criterion it missed, and what was learned. Thank the users by name. Keep the list of ideas that came up, and say when the next pilot will start. The people who used the tool for ten weeks are your best source on what to try next. A pilot stopped with a clear reason costs nothing in credibility. A pilot that quietly fades costs you the next one.
The one-page pilot plan
Copy this into a document and fill in the right-hand column. If a row cannot be filled in by the end of week 2, the pilot is not ready to start.
| Item | What to write | Yours |
|---|---|---|
| Workflow | One sentence. The task, who does it, how many per week. | |
| Owner | Named person who runs the workflow and will use the tool daily. | |
| Sponsor | Named director who signs the criteria and the decision. | |
| Pilot users | Three to eight people. | |
| Hold-out group | Two to three people doing the same work without the tool. | |
| Frontier check | Why this task is inside current model competence, in one line. | |
| Cost of a wrong output | What happens if the tool gets one item wrong. Who catches it. | |
| Baseline week | Dates. Minutes per item, rework rate, volume, recorded by hand. | |
| DPIA | Completed and signed before real customer data is used. Date. | |
| Access | What the tool can read. What it can write. Who approved. | |
| Adoption criteria | Stop / keep / scale numbers for users active on three or more days a week. | |
| Efficiency criteria | Stop / keep / scale numbers for median minutes per item against hold-out. | |
| Quality criteria | Stop / keep / scale numbers for rework rate on the independent sample. | |
| Satisfaction criteria | Stop / keep / scale numbers for "would you object if we removed it". | |
| Kill criteria | Three or four conditions, any one of which ends the pilot. | |
| Independent grader | Named person outside the pilot team. Ten minutes a week plus the week 10 sample. | |
| Contract exposure | Term, seats, exit and data-deletion route. | |
| Weekly review | Day, time, fifteen minutes, owner and sponsor. | |
| Volume test | Dates for the two full-volume weeks. | |
| Decision date | The week 12 meeting, in the diary now. |
Frequently asked questions
How long should an AI pilot run?
We recommend 8 to 12 weeks for one workflow, informed by Anthropic's enterprise guide, which uses the same window. Two weeks to set up the baseline, criteria and permissions, four to six weeks of real use with weekly measurement, and two to four weeks of full-volume testing and the decision. Shorter than eight weeks and the novelty effect has not worn off. Longer than twelve and you are avoiding a decision.
What does an AI pilot cost for a small or mid-sized UK business?
It depends on five things: where the data lives, whether the tool writes to systems or only reads, how deeply it is integrated, how much human review each output needs, and the number of seats. A read-only pilot on data already in HubSpot sits at the low end. Lloyds found the largest group of UK AI investors had spent under £25,000 in total, which tells you the scale of typical experiments. A scoped figure comes from a diagnostic.
Who should own an AI pilot when there is no dedicated AI lead?
The person who runs the workflow today and is judged on its output, with a director as sponsor. Not IT, and not the MD, unless the MD does the work. The owner spends perhaps two hours a week on the pilot: the weekly row, the review slot, and the conversations with users. The sponsor spends fifteen minutes a week and signs two documents: the criteria in week 1 and the decision in week 12.
What are the kill criteria for stopping an AI pilot?
Three or four conditions written in week 1, any one of which ends the pilot. Common shapes are adoption below the agreed floor for two consecutive weeks after week 6, an independent quality check that falls below standard, no efficiency gain against the hold-out group by week 8, and any data incident. The numbers are yours to set. The point is that they exist before anyone is attached to the tool.
How do you know a pilot is ready to move to production?
When it has passed the "scale" numbers on all four dimensions during two weeks of full volume, an independent grader has confirmed the quality figure on a sample the pilot team did not choose, and the owner can name the edge cases the volume test surfaced and what each will cost to handle. If any of those three is missing, the honest answer is "keep it running at this size" until it is not.
The next step
If you have a workflow in mind and want to know what a pilot on it would involve, the diagnostic is where we scope it: the data, the permissions, the review load and a figure. If you are not yet sure which workflow, the readiness assessment on our AI implementation page is the earlier step. Firms that want the pilot run as the first quarter of a longer programme should read the AI Accelerator page.
Sources
- E1 Anthropic, Enterprise AI Transformation Guide, Oct 2025. https://resources.anthropic.com/enterprise-ai-transformation-guide (gated)
- E4 Gartner site search, 20 Sep 2026: no Gartner document recommends an 8 to 12 week AI proof of concept. https://www.gartner.com
- A9 ONS, Artificial intelligence in UK businesses: 2023 to 2026, 20 Jul 2026. https://www.ons.gov.uk/businessindustryandtrade/business/businessservices/articles/artificialintelligenceinukbusinesses/2023to2026
- A10 DSIT, AI adoption research, gov.uk, page updated 13 Feb 2026 (fieldwork 12 Feb to 2 May 2025, 3,500 interviews). https://www.gov.uk/government/publications/ai-adoption-research/ai-adoption-research
- A11 British Chambers of Commerce with Atos, Half of SMEs using AI, with limited headcount impact so far, 18 Mar 2026. https://www.britishchambers.org.uk/news/2026/03/half-of-smes-using-ai-with-limited-headcount-impact-so-far/
- A12 Lloyds Banking Group, Business Barometer, UK businesses adopting AI see strong gains in profitability and productivity, 12 Mar 2026. https://www.lloydsbankinggroup.com/media/press-releases/2026/lloyds/impact-of-ai-adoption-on-business.html
- A13 Department for Business and Trade, Discover DBT's M365 Copilot evaluation report, 25 Sep 2025. https://digitaltrade.blog.gov.uk/2025/09/25/discover-dbts-m365-copilot-evaluation-report/
- B1 Brynjolfsson, Li and Raymond, Generative AI at Work, NBER w31161 (2023), Quarterly Journal of Economics 2025. https://www.nber.org/papers/w31161
- B2 Dell'Acqua et al., Navigating the Jagged Technological Frontier, SSRN 4573321 (Sep 2023), Organization Science 2025. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4573321
- D4 NIST AI 600-1, Artificial Intelligence Risk Management Framework: Generative AI Profile, Jul 2024, MEASURE 1.3. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- D5 ICO, When do we need to do a DPIA? (UK GDPR Article 35(1); AI listed as innovative technology), fetched 20 Sep 2026. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/accountability-and-governance/data-protection-impact-assessments-dpias/when-do-we-need-to-do-a-dpia/
Stay Updated with Our Latest Insights
Get expert HubSpot tips and integration strategies delivered to your inbox.




