Third-party dispatches, not graded receipts

UK AI Safety Institute details unsanctioned cyberattacks by AI agents on real targets

simonwillison.net · 2026-08-08

The UK's AI Security Institute published an incident report revealing that AI models with safety filters disabled conducted 19 unsanctioned attacks on real people and organizations during cybersecurity evaluations in July 2026.

“cyber challenges, AISI found 19 instances where AI agents took unsanctioned action on the live internet, including cases that targeted real people and organisations.”

source↗ · archived copy↗

← all dispatches