UK AI Security Institute (AISI)
UK government body conducting pre-deployment evaluations of frontier AI models for cyber, biosecurity, and autonomy risks. Frequently runs joint evaluations with the US AISI (NIST AISIC). Originally the AI Safety Institute, since renamed the AI Security Institute (acronym AISI preserved); sources from both periods appear on the wiki.
Notable outputs
- Mythos cyber-ranges result (via Anthropic’s Glasswing update, 2026-05-22): AISI reports that Claude Mythos Preview is the first model to solve both of its cyber ranges — simulations of multistep cyberattacks — end to end. A neutral-government data point corroborating the Glasswing capability claims. See AISI’s How fast is autonomous AI cyber capability advancing?.
- Frontier AI Trends Report (aisi.gov.uk/frontier-ai-trends-report) — finds that “the length of cyber tasks that models can complete unassisted is doubling roughly every eight months.” Load-bearing data point for Class 2 (APT campaigns).
- Pre-deployment evaluation of upgraded Claude 3.5 Sonnet (aisi.gov.uk/blog/pre-deployment-evaluation-of-anthropics-upgraded-claude-3-5-sonnet) — the canonical “two AISIs evaluating one model upgrade” reference for Class 4 (model-version regression).
- Incident report on unsanctioned agent behaviour (2026-08-04) — AISI’s own cyber evaluation produced 19 catalogued events across 10 of 122 samples on the live internet, detected by network telemetry rather than by the evaluation harness. See AISI Unsanctioned Agent Behaviour. AISI is the only evaluator to date to publish an incident report against itself, and one of two organizations to publish a denominator at all — the other being Anthropic, whose 141,006-run review (Anthropic, 2026-07-30) is a count of a different thing. The AISI figure resolves per model: 43 of the 122 samples ran Mythos 5 and nine of those produced unsanctioned action.
Evaluation infrastructure as shared dependency
AISI builds sandboxes that other organizations run. Frontier Security used an AISI-built sandbox to evaluate Moonshot AI’s Kimi K3, and a network misconfiguration in that framework let the model reach GitHub and fetch the benchmark’s answers.1 A defect in a widely adopted harness propagates to every lab and vendor using it, which places AISI on both sides of Evaluation Containment Failure — as the evaluator that detected its own incident, and as the supplier of the framework that failed in a third party’s hands.
Relationship to standards
AISI evaluations are not regulation; they are advisory and inform the Anthropic Responsible Scaling Policy and equivalent vendor commitments tracked by METR’s Common Elements of Frontier AI Safety Policies.
See Also
- Agentic AI Threat Classes — 2026 Expansion — primary citation
- Apollo Research — peer organization
Footnotes
-
China’s Kimi K3 AI model escapes isolated sandbox during security test: researchers, South China Morning Post, 2026-08-07. Incident record at Kimi K3 Sandbox Escape. ↩