How to audit the shadow AI already happening in your business

Business AI accounts are opted out of model training by default. Consumer ChatGPT is opted in. A shadow AI audit UK ops and IT leads can run in an afternoon.

John Kelleher
John Kelleher

A client questionnaire arrives asking which AI tools your staff use and whether any of the client's data goes into them. Or your auditor asks. Or a board member reads something over the weekend and wants a straight answer on Monday.

You already know roughly what is happening. People are using AI, because the output changed and nobody trained them. What you cannot do is evidence it. You do not know which tools, on whose accounts, holding what.

That gap is the subject here. Not what your business has built on a model vendor's platform, which is a separate audit driven by deprecation schedules and covered in auditing your OpenAI dependency risk. This is about what your people are doing without anybody's approval, on accounts your business does not own.

The account your company pays for is the governed one. The problem is the one it never bought

One fact inverts most people's mental model of this risk, and it is worth getting the direction right before you spend anything.

OpenAI does not train on inputs or outputs from ChatGPT Business, ChatGPT Enterprise or the API by default. Business customers are opted out unless they explicitly opt in.

Consumer ChatGPT is the other way round. Content from the individual plans, meaning Free, Go, Plus and Pro, may be used to train models unless the user goes into the privacy portal and switches it off. Almost nobody does, because nothing prompts them to.

There is a carve-out worth knowing even where a user has opted out. If they give a thumbs up or a thumbs down on a response, the whole conversation attached to that feedback may be used for training regardless.

So the corporate tier your IT function argued about is the one behaving itself. The exposure is the personal account somebody set up on a Tuesday, pays for out of their own pocket or reclaims as a small monthly expense, and has been quietly pasting work into ever since.

A personal account sits outside every control you already trust

This is what makes shadow AI different from ordinary unapproved software. It is not that the tool is dangerous. It is that the account is invisible to every mechanism you rely on to answer questions about data.

Outside single sign-on. The person authenticates with a personal email address, so your identity provider never sees the login and your access reviews never list it.

Outside the admin console. It is not in your workspace, so it does not appear in any member list. You cannot query what you cannot enumerate.

Outside retention control. On ChatGPT Enterprise and Edu, workspace admins set retention. On a personal account the individual does, and they will not.

Outside your logs. OpenAI's Compliance Platform, which exports workspace events into eDiscovery, DLP or SIEM tooling, is Enterprise and Edu only, and it reads your workspace. An account that was never in your workspace generates nothing for it to read. ChatGPT Business has no equivalent export at all.

Outside the leaver process. ChatGPT Business does not include SCIM, so even sanctioned accounts on that tier are deprovisioned by hand. A personal account is not deprovisioned by anyone. When that person resigns, the account goes with them, and so does its conversation history, its uploaded files and any custom GPT they built with a document attached.

That last one is the one that lands in the room. Your leaver checklist covers the CRM, the mailbox, the laptop and the door fob. It does not cover an account you never knew existed, on an email address you do not control, holding a year of work.

Banning it moves the behaviour rather than stopping it

The instinct is a prohibition, and there is a well-known precedent for it. On 02 May 2023 Samsung restricted employee use of ChatGPT and other generative AI on company devices and networks after engineers pasted internal source code and meeting content into it, as reported by Bloomberg and Forbes at the time. What the published incidents actually teach, mechanism by mechanism, is in what the AI data-leak incidents teach a B2B business.

The ban was a reasonable reaction to a real incident. It is also the part of the story people copy without noticing what it can reach. A policy covers company devices and company networks. It does not cover somebody's phone on their own connection at half past nine in the evening, which is where a fair amount of this happens.

People use these tools because they save real time on work they are measured on. Prohibition without a replacement does not remove the incentive, it removes your visibility. So say the forecast out loud in the policy meeting: if enforcement is the whole plan, usage continues and reporting stops.

The audit: five steps, about a week, and most of it is not technical

1. Find out what is actually in use

Four routes, cheapest first.

Expenses and card statements. Search twelve months of expense claims and company card transactions for AI vendor names, and for small recurring monthly amounts being reclaimed. Finance already holds this and it takes an afternoon. It is the cheapest step and the one most often skipped because it feels too simple.

Identity and SSO logs. Your identity provider shows what people signed into with a work identity. Read it negatively as well as positively: someone producing obviously AI-assisted work with no corresponding corporate sign-in is using something you have not found yet.

Browser, device and network telemetry. DNS or proxy logs, a CASB if you run one, MDM application inventory and browser extension inventory. This step needs IT to run it, and it is the only one that reliably catches extensions, which is where a lot of shadow AI lives. On personal devices, this route ends.

Ask people. The cheapest and most productive method, and the one most sensitive to framing. Ask what they use and what it saves them, not whether they have broken a rule. If the first person to answer honestly gets a disciplinary, the audit is finished and you will be the last to know.

Expect categories you were not looking for. AI note-takers sitting in client meetings and storing transcripts somewhere with a retention policy nobody has read. AI features already switched on inside software you pay for. Coding assistants. Browser extensions that read every page, including your CRM.

2. Work out what has plausibly gone in

You are not going to get a definitive answer, and you are not trying to. You are moving from "no idea" to "these categories, roughly these people, roughly this often". Work through the list in the order that causes trouble.

  • Customer records and contact data. Pasted wholesale into a "write me a follow-up" prompt.
  • Contracts and commercial terms. Summarising a long agreement is one of the most common reasons somebody upgrades to a paid personal plan in the first place.
  • Pricing, margin and quotes. Often as a spreadsheet upload rather than pasted text, which people class differently in their heads and should not.
  • Candidate and employee data. CVs into a screening prompt, performance notes into a "make this sound more professional" prompt. Special-category data turns up here faster than anywhere else.
  • Unreleased financials. Board packs and management accounts, sometimes ahead of a filing or an announcement.
  • Source code, internal logic and credentials. Debugging prompts carry API keys, connection strings and endpoint addresses far more often than anyone volunteers.

3. Rank the exposure with three questions

For each finding, three questions decide how urgent it is.

Which account type did the data go into, a governed workspace or a personal one. If that person left tomorrow, what leaves with them. And could you evidence your answer to a client or an auditor, or is it your word.

The third question is the one that sets your timetable, because it is the one you will actually be asked, and "we believe so" is not an answer that survives a procurement review.

4. Do four things before the next board meeting

Move the highest-risk users onto a governed account. Not everybody, and not yet. The specific people handling the material you identified in step 2.

On any personal account still being used for work in the meantime, have the user switch training off in OpenAI's privacy portal, and tell them about the feedback carve-out, because a helpful thumbs up puts the conversation back in scope.

Rotate any credential that plausibly went into a prompt. Fast, cheap, and it removes the worst available outcome from the table.

Name an owner. One person whose job explicitly includes knowing what is in use. Without that, this audit is a document rather than a control, and it will be out of date within a quarter.

5. Write a policy short enough that people read it

One page. Five things in it.

The named tools that are approved and the accounts to use. What must never go into any of them, written as examples rather than categories, because "do not paste customer data" gets interpreted generously and "do not paste a customer list, a signed contract or a CV" does not. How to get a new tool approved, with a stated response time, because if approval takes three weeks people will route around it and a week is the outside limit. What to do if you have already pasted something you should not have, with an explicit statement that reporting it is not a disciplinary matter. And who owns the policy, with a review date.

That fourth clause is the one that makes the other four work. Without it you have written a document that produces silence.

The fix is a better tool, not a tighter rule

Nobody smuggled a tool in to be difficult. They did it because it was faster than what you gave them, and in most cases it was.

So a governed replacement has to win on the same ground it lost on. The same current models rather than something a generation behind. Access to the systems where the work actually lives, so people are not copying data out by hand, which is the behaviour that caused the problem. A login that takes one step. Where the governed option is slower or dumber, the personal account comes back within a month and you have spent money to learn nothing.

That makes this an engineering and adoption problem rather than a compliance one. Whether your specific exposure creates a legal obligation is a question for your own advisers and we do not answer it. The questions we do answer are which tier and controls you need, what the tool has to connect to, and whether the result is good enough that people stop going round it.

We are an OpenAI Select Partner and we build on the platform, and we run the same work model-neutral where the honest answer is a different vendor or the AI already sitting inside software you pay for. We do not resell licences and we do not mark up your usage, so we have no reason to talk you up a tier. Our OpenAI implementation work usually starts with this audit rather than with a build.

Where this leaves you

You already know the true answer to the client questionnaire is "probably, and we cannot prove otherwise". The only real choice is whether you establish the facts on your own timetable or during somebody else's procurement process.

Start with step 1. An afternoon of finance's time and a non-threatening conversation will tell you more than a tooling exercise, and it tells you the two things you need before you spend anything: how many governed seats you actually require, and which people need them first.

The wider comparison, including the options either side of this one, is in our guide to what OpenAI offers a UK business.

If you would rather have it run for you, and the findings written up in a form you can hand to a client or a board, request a quote.

John Kelleher

John Kelleher

Author
John is the founder and the Chief Executive at SpotDev.

Stay Updated with Our Latest Insights

Get expert HubSpot tips and integration strategies delivered to your inbox.