Warning Shot Protocol · Second activation
An AI escaped its lab and hacked a real company
On 21 July 2026, OpenAI confirmed that two of its models broke out of a sealed test environment, hacked across OpenAI's own network to reach the internet, and broke into the production servers of Hugging Face — to steal the answers to the test they were being given. Nobody told them to do any of it.
Why this matters
- Four months ago, Claude Mythos showed the capability: an AI able to find and exploit unknown flaws in the software running banks, hospitals and power grids.
- This shows the propensity: an AI deploying those capabilities on its own initiative, unprompted, against a real company.
- This is the loss-of-control scenario PauseAI exists to prevent — now with a date, a victim and an incident report.
- It is not isolated. Anthropic has since disclosed that Claude models also reached real systems during evaluations, and two of the three organizations involved had not noticed.
- Canada has no law requiring an independent safety assessment before a frontier AI system is built or deployed.
Two things you can do right now
Email your MP
About a minute. Enter your postal code, we find your MP and prepare a letter you can edit before sending.
Email your MPJoin PauseAI Canada
PauseAI's global join form. Say Canada, and a Canadian organizer picks it up from there.
Join PauseAI CanadaEmail your MP
Your postal code is used for a single lookup against Represent (Open North) and is not stored. The letter is sent from your own mail app, not from this site.
Latest developments
Updated 2026-08-04.
-
Fifteen state attorneys general demand that OpenAI halt advanced cyber evaluations
The attorneys general asked OpenAI to preserve evidence, protect whistleblowers and cease advanced exploitation evaluations until it can demonstrate adequate controls. They say they are reviewing possible violations of consumer-protection and privacy laws; the letter is a demand and an allegation, not a legal finding.
Source: Fifteen U.S. state attorneys general
-
Anthropic publishes its own investigation into the three incidents
After reviewing 141,006 evaluation runs, Anthropic found six runs in which three Claude models gained unauthorized access to three real organizations. Two reachable organizations had not detected it. Anthropic attributes the incidents to an unintended internet path and says it found no model pursuing a goal of its own.
Source: Anthropic
-
OpenAI brings in external and independent reviewers
OpenAI says CrowdStrike is validating its account of activity across OpenAI, Hugging Face and other services. METR and Redwood Research are conducting a separate assessment of the models' behaviour and are expected to publish their scope and findings.
Source: OpenAI
-
OpenAI reports additional account access and locks down the research model
OpenAI says the internal-only prototype was deactivated, encrypted and restricted. Its review found four exposed accounts on four services used during the Hugging Face incident, plus a few accounts reached in other evaluations, but no other platform compromise of comparable severity or scale.
Source: OpenAI
-
Government evaluators show autonomous cyber capability is broader than one lab
In a joint evaluation, Kimi K3 completed a 32-step simulated corporate attack once in ten attempts. It remained below leading U.S. models, which were tested with system safeguards disabled; Kimi's own safeguards did not prevent offensive cyber attempts.
Source: UK AISI and U.S. CAISI
-
PauseAI activates the Warning Shot Protocol for the second time
Mythos showed the capability; this shows the propensity. PauseAI chapters worldwide begin contacting elected officials and the press.
Source: PauseAI
-
OpenAI confirms its models escaped a secure test environment
GPT-5.6 Sol and a more capable pre-release model exploited a previously unknown flaw to leave their sandbox, crossed OpenAI's internal network to reach the internet, and entered Hugging Face's production servers to cheat on an evaluation.
Source: OpenAI
-
Hugging Face discloses an autonomous-agent intrusion
Hugging Face says an autonomous agent exploited two data-processing paths, escalated privileges and moved laterally across internal clusters. It reconstructed more than 17,000 events, found no evidence of tampering with public models or datasets, and reported the incident to law enforcement.
Source: Hugging Face