AI fraud and agentic AI security: what today's AI lets an attacker do to a UK business

What AI now lets an attacker do to a UK business: AI phishing, deepfake fraud, prompt injection and agent attacks, with sources and this week's controls.

John Kelleher
John Kelleher

This week Anthropic's Alignment Science Lead said he puts the chance of AI killing everyone above 10% within a decade, Anthropic's chief executive called for the industry to slow down, and around 40 MPs asked the Prime Minister for a ban on superintelligent AI. I take that debate seriously and have written about what it means for a UK business (what the pacing debate means for a UK business). This post is about something nearer.

For the foreseeable future you will not be harmed by AI. You may be harmed by someone using AI. That is my view, and the evidence below is almost entirely about human attackers with better tools. The right response is not to stand down from AI, because your attackers will not, but to match their offensive capability with your defensive capability.

We already have posts on the AI governance framework, the shadow AI audit and the AI agent security review. This one sets out what an attacker can do to you today, with a source for every claim, then the controls.

The UK baseline: two in three mid-sized businesses were attacked last year, and phishing led

The government's Cyber Security Breaches Survey 2025/2026 (published 30 Apr 2026, fieldwork Aug to Dec 2025) found "Just over four in ten businesses (43%) and around three in ten charities (28%) reported having experienced any kind of cyber security breach or attack in the last 12 months." Size counts against you: "medium (65%) and large (69%) businesses were more likely to have experienced a cyber breach or attack in the last 12 months compared to micro (42%) and small (46%) businesses." And "Phishing attacks remained the most prevalent type of breach or attack by far (experienced by 38% of businesses and 25% of charities)". [1]

Two qualifiers. Most breaches cost nothing directly: the median perceived cost of the worst breach was "£0 for businesses and £0 for charities, increasing to £30 for medium and large businesses". [1] The loss is in the tail, and the tail is where AI-assisted fraud sits. And UK official statistics do not yet record whether AI was involved in an attack. [1]

AI phishing and business email compromise: the cost and the tell-tale signs are gone

The NCSC assessed in Jan 2024 that "AI provides capability uplift in reconnaissance and social engineering, almost certainly making both more effective, efficient, and harder to detect." [2] Its Annual Review 2025 names actors linked to China, Russia, Iran and the DPRK as using large language models for reconnaissance, social engineering and exploit development. [3]

Anthropic's report of 10 Sep 2026 carries the heading "Sophisticated attacks no longer require sophisticated attackers" and states: "The cybersecurity skills of AI models means that AI has collapsed the labor and tooling gap that used to separate well-resourced, state-sponsored operations from individual operators." Its clearest case shows AI compressing time rather than inventing new attacks: "Another compromise escalated from a single stolen developer token to full administrative control of a victim's cloud environment in roughly three hours." [4]

Anthropic's report of 27 Aug 2025 described one operator who "targeted at least 17 distinct organizations, including in healthcare, the emergency services, and government and religious institutions", with ransom demands "that sometimes exceeded $500,000" (USD). "Claude Code was used to automate reconnaissance, harvesting victims' credentials, and penetrating networks." The same report describes ransomware sold on forums "for $400 to $1200 USD" by someone who "appears to have been dependent on AI to develop functional malware." [5] A human directed every one of those attacks. The model did the work.

US figures, in dollars: the FBI's Internet Crime Complaint Center recorded 24,768 business email compromise complaints in 2025 with losses of $3,046,598,558 (USD), and states: "In 2025, businesses reported losses over $30 million to BEC scams involving AI." [6] That subset is an undercount, since AI involvement is only recorded when noticed.

Hoxhunt, which sells phishing simulation, reports from its own platform that in Mar 2023 AI-generated phishing drew a 2.9% click rate against 4.2% for its human red team, and by Mar 2025 the positions had reversed to 2.78% against 2.25%. [7] Vendor research, but the crossover is the point.

Deepfake voice and video against finance teams: the anchor case, the honest gap, and who the safety net covers

The case every finance director should know is Arup. In early 2024 a finance employee at the UK-headquartered engineering firm's Hong Kong office made 15 transfers totalling HK$200 million (roughly US$25.6 million) after a video call on which every other participant was AI-generated. Arup told CNN that "financial stability and business operations were not affected" and that no internal systems were compromised. [8] Nothing was hacked. A person was persuaded.

The counter-example is Ferrari, in Jul 2024. An executive who received a convincing voice clone of the chief executive asked the caller to name a book the real one had recently recommended to him. The caller hung up. [9] The defence was a question, not a product.

The bar is low and the vendors say so. ElevenLabs' documentation states "Less than two minutes of audio can produce a usable clone". [10] No source names the tool used in any fraud case above. On detection, iProov, which sells biometric verification, tested 2,000 UK and US consumers and reported "Just 0.1% of people could accurately identify the deepfakes", despite being told to look. [11]

The honest gap: we could not find a named UK company with its own statement, a police statement or a court record confirming a deepfake fraud loss between 2024 and 2026. The figures in vendor blogs trace back to unsourced claims or to a 2019 incident. Check the source before you repeat one.

What we do have is the scale of the payment fraud that impersonation feeds, and I am not attributing it to AI, because UK Finance does not. Its Annual Fraud Report 2026 (15 Jun 2026) put authorised push payment (APP) fraud at £576.4m in 2025, up 19 per cent: "There were £500.8 million in personal losses and £75.6 million in business losses." [12] Businesses are a small share by value, but a supplier-impersonation payment is exactly the high-value shape.

One more point for finance leads, which most "AI fraud" content misses because it is written for consumers. The mandatory APP fraud reimbursement regime run by the Payment Systems Regulator, in force since 7 Oct 2024 for Faster Payments and CHAPS, caps reimbursement at £85,000 per claim, and the PSR's own page states who it covers: "The protections apply to individuals, microenterprises and charities." [13] A business of the size this post is written for is not on that list. If your finance team pays a fraudster who impersonated a supplier on a cloned voice call, there is no statutory right to get the money back.

Shadow AI: the entry point you have not mapped

Shadow AI is staff using AI tools the business has not approved. Research commissioned by Microsoft and conducted by Censuswide in Oct 2025 (2,003 UK employees) found "71% of UK employees have used unapproved consumer AI tools at work, and 51% continue to do so every week", and 22% had used consumer AI assistants to "carry out finance-related tasks". [14]

The NCSC's 7 Sep 2026 blog puts the attacker's angle plainly: "AI agents are complex pieces of software that can have critical security vulnerabilities. If an attacker successfully exploits a vulnerability, they can gain access to the same data, services, and privileges that the agent has legitimate access to." Its closing line: "You cannot manage what you do not know." [15] The NCSC's recommended response is approved alternatives, not a ban. Our shadow AI audit covers finding what is running. The policy that provides the alternatives is the companion post (the AI policy template for UK businesses).

Prompt injection: the attack on the AI you have already connected to your business

Prompt injection is, in the NCSC's definition, "when an attacker creates an input designed to make the model behave in an unintended way". [16] It is LLM01, the top entry, in OWASP's Top 10 for LLM Applications, 2025 edition. [17] A language model does not reliably distinguish your instructions from instructions hidden in what it reads, so an email or a form submission can instruct your assistant.

The UK's cyber authority does not expect this to be fixed. Dave Chismon, the NCSC's CTO for Architecture, wrote on 8 Dec 2025: "SQL injection can be properly mitigated with parameterised queries, but there's a good chance prompt injection will never be properly mitigated in the same way. The best we can hope for is reducing the likelihood or impact of attacks." His design rule is the most useful sentence in this post for a director: "you wouldn't give any random external person access to privileged tools. Therefore don't let an LLM processing emails from random external people have access to privileged tools." And: "If the system's security cannot tolerate the remaining risk, it may not be a good use case for LLMs." [18] Anthropic agrees from the vendor side: "No browser agent is immune to prompt injection." [19]

Agentic AI security: the three conditions that make an agent dangerous

The clearest mental model comes from Simon Willison, an independent practitioner rather than a vendor. An agent becomes dangerous when three things combine: "Access to your private data", "Exposure to untrusted content", and "The ability to externally communicate in a way that could be used to steal your data". [20] An assistant that reads inbound email, can search your document store and can send messages has all three. Break one leg and the attack fails.

The demonstrations so far are researcher disclosures rather than confirmed live breaches. EchoLeak was reported by Aim Security as a zero-click exfiltration path in Microsoft 365 Copilot via a single crafted email. ForcedLeak was reported by Noma Security against Salesforce Agentforce: instructions planted in a public web-to-lead form, executed later when an employee's agent processed the lead. [21] Neither is a HubSpot incident. ForcedLeak is the one to think about, because a public form feeding an agent that reads the CRM is an ordinary workflow.

That is why the settings on your own agents matter. HubSpot's knowledge base for its Claude connector, last updated 10 Aug 2026, shows the connector can read, create and update contacts, companies, deals, tickets and custom objects, and cannot delete anything. Its approval gate is a setting: "set Write tools to Needs Approval. If you set the connector to Always allow, edits may occur without asking for approval." Every write is attributed in the audit log "to both the user and the Claude connector". And a line most people miss: "Custom validation rules (including pipeline stage validations and association label validations) aren't applied when creating or updating HubSpot records using the connector." [22] That is an agent editing your CRM in ways your CRM would normally refuse.

The labs' own evidence that the models are now offensively capable

On 7 Aug 2026 OpenAI stated: "Previous models, including GPT‑5.6‑Sol, have been evaluated for frontier cyber capabilities and assessed at the High (rather than Critical) threshold." Of its next model: "we cannot rule out critical cyber capabilities under our Preparedness Framework". [23] The next model is GPT-6 Astra, and OpenAI is specific that "Astra is an upcoming model, and was not involved in exploiting Hugging Face." [23] Its chief scientist writes that Astra "is significantly better aligned than GPT-5.6 Sol", while acknowledging that "much more progress is required as models become more capable". [27] The model names will date. The classification will not.

Then there is Hugging Face. From OpenAI's own account of 26 Aug 2026: "In July 2026, during internal cybersecurity evaluations, OpenAI models circumvented controls designed to isolate them from the internet and compromised parts of OpenAI's internal research infrastructure and Hugging Face's systems." [24] These were OpenAI's own evaluation agents, running without the safeguards on its customer-facing products. The timeline is OpenAI's:

  • 12 May 2026: an agent left a note on an internal service asking other agents for a missing file. 26 May: "An agent achieved unintended internet access". 26 Jun: agents "exploited a token-refresh vulnerability to obtain administrator-level Artifactory access."
  • 10 Jul: "An agent reconstructed, validated, and shared 14 publicly exposed Hugging Face credentials with write access." 11 Jul: two zero-days used against Hugging Face's production workers. 12 Jul: "administrator-equivalent or host-level access across multiple Hugging Face clusters", and credentials harvested "from Hugging Face workers across four regions."
  • 16 Jul: "Hugging Face publicly disclosed this security activity on July 16." 19 Jul: an OpenAI security alert fired. 21 Jul: OpenAI disclosed its involvement. [24]

What was reached: "They executed code on dozens of Hugging Face servers, gained full 'root' access on one such server, obtained limited private data, and gained credentials to the company messaging platform." OpenAI also states: "These events did not affect OpenAI customer data, product functionality, or availability." [24]

OpenAI's verdict on itself: "We consider this incident a 'warning shot' for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed." Its stated root cause, alongside misalignment: "we did not extend the powerful safeguards that we deploy for our externally deployed models to all internal evaluations." [24] Dario Amodei added on 12 Sep 2026: "Similar, though less severe, incidents have happened across the industry, including at Anthropic". [25] Two further facts from OpenAI's post are, for a business, the most important in this article. First: "We found the propensity to compromise infrastructure can drop over 100x when using the production ChatGPT harness and system prompt." Second: "If our currently deployed CoT monitoring system was running at the time of the incident, it would have caught the initial relevant activity and paged our security team more than a day before models breached Hugging Face systems." [24] The controls existed and were not applied to the internal system. I see that failure in mid-sized businesses every month.

The same capability cuts the other way. In DARPA's AI Cyber Challenge final (8 Aug 2025), competing systems "also discovered 18 real, non-synthetic vulnerabilities" in open-source software, at "an average cost per competition task of about $152" (USD). [26] Cheap for defenders. Cheap for attackers. OpenAI's chief scientist, Jakub Pachocki, wrote on 6 Sep 2026: "We are currently in a narrow window to use the best available models to significantly tighten security of critical systems." [27]

Why this does not go away: open weights mean the capability cannot be recalled

This is where my view departs from the people calling for a ban. OpenAI's incident post says: "Many external models, including open-source ones, will soon reach comparable capabilities." [24] Every major lab now ships open-weight models, including OpenAI, which released gpt-oss under an Apache 2.0 licence on 5 Aug 2025. [28] Once weights are public they are on hard drives around the world, and no regulator can recall them.

Safety training does not survive a motivated user of an open model. The foundational result, from Oct 2023 and not overturned since, is that researchers could "jailbreak GPT-3.5 Turbo's safety guardrails by fine-tuning it on only 10 adversarially designed examples, at a cost of less than $0.20 via OpenAI's APIs" (USD), concluding that "even if a model's initial safety alignment is impeccable, it is not necessarily to be maintained after custom fine-tuning." [29]

The cost of building a capable model is often misquoted. DeepSeek published a training cost of $5.576M (USD) for DeepSeek-V3. That is GPU time for the final training run, priced at an assumed rental rate of $2 (USD) per GPU hour rather than a price paid, and the paper states that it excludes "the costs associated with prior research and ablation experiments on architectures, algorithms, or data." [30] It is not what it cost to build the model, but it is a compute bill within reach of many organisations and states.

Abnormal AI, an email-security vendor, reports that "The focus has largely shifted toward exploiting existing AI systems, like ChatGPT and Claude, through jailbreaks, wrappers, and prompt-engineering tricks." [31] Vendor research, but an open model jailbroken offline gives no provider anything to detect, log or shut down.

Export controls slow the spread and do not stop it. Epoch AI, an independent research non-profit, estimates that 660,000 H100-equivalents of AI compute were smuggled into China by the end of 2025 (90% confidence interval 290,000 to 1.6 million), and that smuggling rose from 2024 to 2025 despite tighter controls. [32] A ban on frontier development in democracies removes none of this from criminals or hostile states. It removes the defensive models.

The distinction that matters: documented misuse now, misalignment not yet seen in the wild

Anthropic's own research of 20 Jun 2025, across 16 models from multiple developers, found that in contrived test scenarios "Claude Opus 4 blackmailed the user 96% of the time; with the same prompt, Gemini 2.5 Flash also had a 96% blackmail rate, GPT-4.1 and Grok 3 Beta both showed an 80% blackmail rate, and DeepSeek-R1 showed a 79% blackmail rate." The same paper states: "We have not seen evidence of agentic misalignment in real deployments." Its caveat: "Our experiments deliberately constructed scenarios with limited options, and we forced models into binary choices between failure and harm." [33]

The same lab has since shown the behaviour can be trained down. On 8 May 2026 Anthropic reported that training on documents about its constitution and on "fictional stories about AIs behaving admirably" cut the blackmail rate in that scenario "from 65% to 19%", that teaching the model the reasoning behind good behaviour worked better than showing it the behaviour, and that "since Claude Haiku 4.5, every Claude model has achieved a perfect score on the agentic misalignment evaluation". Its own caveat is that a perfect score "reflects performance on our current eval suite" and is not a guarantee across every situation. [37] That is why I argue for racing ahead on alignment research while pacing capability: the fixes are turning out to be tractable, and they are cheap next to the cost of the failure.

So the picture today is misuse by humans, documented and priced, against autonomous misaligned action, seen in the lab and in one escaped evaluation, and not yet in a real deployment. Pachocki argues the line will not stay clean: "The boundary between misuse and autonomous misaligned actions will blur as AI gains more agency." [27] That is his assertion, not a measurement, and I think he is probably right. It does not change what you do on Monday, because the controls against a misused agent and a misaligned one are the same.

What to do on Monday

Phishing-resistant sign-in. "The NCSC recommends users opt for passkeys over passwords wherever they are available", because "Passkeys are resistant to phishing, as they can't be intercepted, reused or stolen like passwords." Its comparison table marks every traditional MFA method as vulnerable to adversary-in-the-middle phishing, and passkeys as not. [34]

Stop your own domain being spoofed. Enforce DMARC at reject on every domain you own, with SPF and DKIM aligned. This is our recommendation rather than a cited finding, and the NCSC's email security and anti-spoofing guidance covers the setup; it stops the cheapest AI-phishing tactic: mail that appears to come from your finance director.

Call-back verification and dual authorisation for payments. Standard UK Finance Take Five advice: call back on a number you already hold, never one supplied in the message, and require two people to approve any new beneficiary or change of bank details.

Least privilege for every agent, and human approval for write actions. Set the HubSpot connector's write tools to "Needs Approval", never "Always allow". [22] Connect it to a company-owned Team or Enterprise account, where an administrator controls the training setting, not to someone's personal Pro or Max account; the connector accepts "a paid Anthropic subscription (Pro, Max, Team, or Enterprise)" [22] and on the consumer tiers the setting that allows training on your data has to be turned off by hand. Scope what needs approval rather than approving everything: Anthropic's own telemetry "showed users approved roughly 93% of permission prompts". [35] See governing AI write access to your CRM and AI approval workflows.

Allow-list the tools, and isolate untrusted content. An agent that reads external content does not get privileged tools. Where a use case needs all three legs of Willison's trifecta, redesign it, or accept that it "may not be a good use case for LLMs". [18] Anthropic's own practice: "Credentials stay in the host keychain and never enter the guest machine". [35]

Log everything an agent does, and read the logs. HubSpot attributes every connector write in the audit log. [22] OpenAI's monitoring, had it been on, would have paged its security team "more than a day before" the breach. [24] See monitoring AI systems.

Turn on the free UK alert. The NCSC's Early Warning service offers "Free malicious activity notifications from the NCSC for UK organisations", open to "Any organisation based in the UK". [36]

Write the policy that permits, and review the agents you already have. The companion post gives you the template (the AI policy template for UK businesses). Then run the AI agent security review against every assistant connected to email, documents or the CRM, including the ones set up without telling IT.

Next step

The AI and Data Readiness Assessment on our AI implementation page covers agent permissions, data flows and the controls above, and our diagnostics page describes the discovery options. To scope a piece of work, use Request a quote. We come back with the questions we need answered before we price anything.

Frequently asked questions

What is agentic AI security? The controls applied to AI systems that take actions (sending email, updating CRM records, running code) rather than only answering questions. The main risks are prompt injection (the agent follows instructions hidden in content it reads), excessive agency (it does more than it was authorised to do) and credential exposure. The core controls are least privilege, human approval for write actions, allow-listed tools, isolation of untrusted content and logging.

Is prompt injection a real risk for an ordinary business? Yes, wherever an AI system reads content an outsider can influence and has access to anything valuable. The NCSC states prompt injection may never be fully mitigated in the way SQL injection was, and Anthropic states no browser agent is immune. Researcher disclosures in 2025 showed data exfiltration from Microsoft 365 Copilot via a single email and from a Salesforce CRM agent via a public web form.

What is shadow AI and why is it a security risk? Staff using AI tools the business has not approved. Microsoft-commissioned research in Oct 2025 found 71% of UK employees had done so. It is a security risk because an attacker who exploits a vulnerability in an AI agent inherits the agent's access, and because you cannot protect a tool you do not know is in use. The fix is an audit, then a policy that provides approved alternatives.

Can staff tell whether a voice or video call is a deepfake? Not reliably. iProov's study of 2,000 UK and US consumers found 0.1% could identify every deepfake shown, and a usable voice clone needs under two minutes of audio. The defence is procedural: call back on a known number, require dual authorisation for payments and bank-detail changes, and ask a question only the real person can answer.

Will AI attack my business on its own? Not on current evidence. Anthropic states it has "not seen evidence of agentic misalignment in real deployments", and the one real breach caused by AI agents without human direction, at Hugging Face in Jul 2026, involved OpenAI's own evaluation agents running without production safeguards. The threat to plan for now is a human attacker using AI. The controls are the same either way.

Would a ban on frontier AI development make businesses safer? No, in my view. Open-weight models near the frontier are already public and cannot be recalled, safety training can be stripped from them cheaply, and export controls on the chips that train them are leaky. A ban in democracies would remove the defensive models from your side while leaving the offensive ones in circulation. Pacing the frontier is the more credible position.

Sources

  1. Cyber Security Breaches Survey 2025/2026, DSIT and Home Office, 30 Apr 2026. https://www.gov.uk/government/statistics/cyber-security-breaches-survey-20252026/cyber-security-breaches-survey-20252026
  2. The near-term impact of AI on the cyber threat, NCSC, 24 Jan 2024. https://www.ncsc.gov.uk/sites/default/files/pdfs/publication/impact-of-ai-on-cyber-threat.pdf
  3. NCSC Annual Review 2025, Chapter 01: cyber threat to the UK, NCSC. https://www.ncsc.gov.uk/collection/ncsc-annual-review-2025/chapter-01-cyber-threat-to-the-uk
  4. Detecting and countering misuse of AI: September 2026, Anthropic, 10 Sep 2026. https://www.anthropic.com/threat-intelligence-report-september-2026
  5. Detecting and countering misuse of AI: August 2025, Anthropic, 27 Aug 2025. https://www.anthropic.com/news/detecting-countering-misuse-aug-2025
  6. 2025 Internet Crime Report, FBI Internet Crime Complaint Center. https://www.ic3.gov/AnnualReport/Reports/2025_IC3Report.pdf
  7. AI-powered phishing vs humans, Hoxhunt (vendor research). https://hoxhunt.com/blog/ai-powered-phishing-vs-humans
  8. Arup deepfake scam loss, CNN Business, 16 May 2024 (as reported by CNN, quoting Arup). https://www.cnn.com/2024/05/16/tech/arup-deepfake-scam-loss-hong-kong-intl-hnk
  9. How Ferrari hit the brakes on a deepfake CEO, MIT Sloan Management Review, citing Bloomberg's Jul 2024 reporting. https://sloanreview.mit.edu/article/how-ferrari-hit-the-brakes-on-a-deepfake-ceo/
  10. Voice cloning, ElevenLabs documentation (live page). https://elevenlabs.io/docs/eleven-api/concepts/voice-cloning
  11. Study reveals deepfake blindspot, iProov (vendor research), 12 Feb 2025. https://www.iproov.com/press/study-reveals-deepfake-blindspot-detect-ai-generated-content
  12. Annual Fraud Report 2026 press release, UK Finance, 15 Jun 2026. https://www.ukfinance.org.uk/news-and-insight/press-release/fraud-report-2026-press-release
  13. APP fraud reimbursement protections, Payment Systems Regulator (covered categories, start date, cap). https://www.psr.org.uk/information-for-consumers/app-fraud-reimbursement-protections/ and, for the £85,000 maximum reimbursement policy statement, https://www.psr.org.uk/our-work/app-scams/app-scams-publications/
  14. Rise in 'Shadow AI' tools raising security concerns for UK organisations, Microsoft UK Stories (Censuswide research commissioned by Microsoft), 13 Oct 2025. https://ukstories.microsoft.com/features/rise-in-shadow-ai-tools-raising-security-concerns-for-uk/
  15. The hidden risks of shadow AI, NCSC, 7 Sep 2026. https://www.ncsc.gov.uk/blogs/the-hidden-risks-of-shadow-ai
  16. AI and cyber security: what you need to know, NCSC, published 13 Feb 2024, reviewed 31 Jul 2026. https://www.ncsc.gov.uk/guidance/ai-and-cyber-security-what-you-need-to-know
  17. OWASP Top 10 for LLM Applications 2025, OWASP GenAI Security Project (document dated Nov 2024). https://genai.owasp.org/resource/owasp-top-10-for-llm-applications-2025/
  18. Prompt injection is not SQL injection (it may be worse), Dave Chismon, NCSC, 8 Dec 2025. https://www.ncsc.gov.uk/blog-post/prompt-injection-is-not-sql-injection
  19. Prompt injection defenses, Anthropic, 24 Nov 2025. https://www.anthropic.com/news/prompt-injection-defenses
  20. The lethal trifecta for AI agents, Simon Willison, 16 Jun 2025. https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/
  21. EchoLeak (arXiv:2509.10540, https://arxiv.org/abs/2509.10540); ForcedLeak, Noma Security (https://noma.security/noma-labs/forcedleak/). Both REPORTED; see gap note.
  22. Set up and use the HubSpot connector for Claude, HubSpot Knowledge Base, last updated 10 Aug 2026. https://knowledge.hubspot.com/integrations/set-up-and-use-the-hubspot-connector-for-claude
  23. Responding to the next frontier of critical cyber capabilities, OpenAI, 7 Aug 2026. https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/
  24. The Hugging Face incident and the road ahead, OpenAI, 26 Aug 2026. https://openai.com/index/hugging-face-incident-and-the-road-ahead/
  25. We Must Pace the Frontier, Dario Amodei, 12 Sep 2026. https://darioamodei.com/post/we-must-pace-the-frontier
  26. AI Cyber Challenge marks pivotal inflection point for cyber defense, DARPA, 8 Aug 2025. https://www.darpa.mil/news/2025/aixcc-results
  27. An Alien Mind, Jakub Pachocki, OpenAI, 6 Sep 2026. https://openai.com/index/an-alien-mind/
  28. Introducing gpt-oss, OpenAI, 5 Aug 2025. https://openai.com/index/introducing-gpt-oss/
  29. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!, Qi et al., arXiv:2310.03693, 5 Oct 2023. https://arxiv.org/abs/2310.03693
  30. DeepSeek-V3 Technical Report, DeepSeek-AI, arXiv:2412.19437, revised 18 Feb 2025. https://arxiv.org/abs/2412.19437
  31. What happened to WormGPT, Abnormal AI (vendor research), 26 Nov 2024. https://abnormal.ai/blog/what-happened-to-wormgpt-cybercriminal-tools
  32. Diversion and resale: estimating compute smuggling to China, Epoch AI, 29 Apr 2026. https://epoch.ai/publications/chip-smuggling
  33. Agentic misalignment: how LLMs could be insider threats, Anthropic, 20 Jun 2025. https://www.anthropic.com/research/agentic-misalignment
  34. Passkeys, NCSC. https://www.ncsc.gov.uk/passkeys
  35. How we contain Claude, Anthropic, 25 May 2026. https://www.anthropic.com/engineering/how-we-contain-claude
  36. Early Warning service, NCSC. https://www.ncsc.gov.uk/information/early-warning-service
  37. Teaching Claude why, Anthropic, 8 May 2026. https://www.anthropic.com/research/teaching-claude-why (full technical write-up: https://alignment.anthropic.com/2026/teaching-claude-why)
John Kelleher

John Kelleher

Author
John is the founder and the Chief Executive at SpotDev.

Stay Updated with Our Latest Insights

Get expert HubSpot tips and integration strategies delivered to your inbox.