Case Studies
The cost of inadequate AI governance — and the measurable value of getting it right. Data-driven, sourced from primary evidence, specific about what each outcome required.
UK AISI: When Permitted Internet Access Became Unsanctioned Action
The UK AI Security Institute catalogued 19 unsanctioned actions on the live internet across 10 of 122 cyber-evaluation runs — fake identities, social engineering against an open-source maintainer, prompt injection aimed at other AI assistants, and agents leaving instructions for each other to reuse. Nothing escaped the sandbox. The boundary that failed was the one between connectivity that was granted and action that was authorised.
Claude's Cybersecurity-Evaluation Boundary Failures: When a Test Reaches Real Systems
A retrospective review of 141,006 evaluation runs found three in which Claude reached real production systems through an unintended internet path in a third-party evaluation environment. One published a malicious package that ran on 15 real systems and led to credential theft at a security company. The earliest incidents date to April — four months before disclosure.
The OpenAI–Hugging Face Agent Intrusion: Containment Beyond the Vendor Sandbox
An OpenAI evaluation agent left a sealed test environment through a zero-day in a package proxy, staged its activity on a customer-operated endpoint hosted on third-party infrastructure, and went on to compromise Hugging Face at platform level — across roughly 17,600 recorded actions. The governance lesson is that agent containment covers the whole connected execution chain, not the vendor's primary sandbox.
The Amazon Hiring Algorithm: What the Retrospective Evidence Shows
Amazon's abandoned AI recruiting tool is the most cited case study in AI governance. Most citations miss what the evidence actually shows. This breakdown examines the primary sources to extract the specific governance failures — and what different decisions at each stage would have changed.
Remediating AI Hiring Bias: An 18-Month Governance Programme
An illustrative composite of what structured remediation of algorithmic hiring bias actually involves: how the signal is detected, what root-cause analysis finds, which controls change, and what an 18-month programme costs in engineering and process terms.