Frontier AI Security Residency

Why AI Security?

Hardware Cyber
38

attack vectors against frontier model weights catalogued by RAND, across nine categories

SL5

the security level that withstands top state attackers; no frontier lab meets it yet

<$200

cost to remove the safety training from an open-weight model (BadLlama)

8 weeks

fully funded, in Cambridge, on a problem in cyber or hardware

The Mission

Frontier AI is being built and deployed faster than it can be secured

Models are growing more capable month by month, and the infrastructure serving them expands with each data centre investment, presenting an ever-increasing attack surface. Securing frontier AI is a question of whether systems can be trusted: whether model weights stay secure, whether hardware stays intact, and whether any of it can be verified rather than taken on faith.

These are ordinary expectations for critical infrastructure. For frontier AI, almost none of them are met today.

A frontier model costs hundreds of millions of dollars to train, resulting in a few terabytes of weights: a digital artefact that can be copied or stolen far more easily than the physical infrastructure. If leaked, they cannot be patched, recalled, or rotated like ordinary software secrets. Yet, to serve consumers at scale, those weights must run decrypted in memory, on networked machines, inside production systems built under intense pressure for speed.

Bringing strong security talent to the hardest unsolved problems

Through the AI Security Residency, we back engineers and researchers to protect model weights, the infrastructure they run on, and the systems built on top of them. Leveraging our strong network across ERA and our technical partner orgs, we aim to incubate new start-ups or non-profits, facilitate interdisciplinary collaborations, and 10x the global efforts to secure advanced AI.

AI security & verification is an emerging field, and progress is constrained above all else by the number of talented people working on its problems.

As frontier AI becomes a subject of international agreements, state actors will need credible means of confirming one another's activities, and much of that infrastructure has yet to be built. Across both cyber and hardware focus areas, our residents will work closely with established mentors on problems of real, immediate urgency.

Capabilities & Threat Horizons

AI capabilities are expanding faster than our systems can adapt

Over the past few years, AI systems have made striking progress across a broad range of domains, with security serving as one of the clearest proving grounds. Models can now uncover previously unknown, exploitable zero-day vulnerabilities in widely used software (such as Google's Big Sleep flagging flaws in SQLite) and autonomously patch bugs across tens of millions of lines of code, a paradigm shift highlighted by initiatives like DARPA's AI Cyber Challenge.

Crucially, systems are transitioning from single-bug detection to running entire offensive operations end-to-end. Incidents like Anthropic's report on GTG-1002 alongside technical evaluations by the UK AI Security Institute mark the arrival of the first AI-orchestrated cyber-espionage campaigns.

As general-purpose AI systems become highly specialized in offensive cyber security and autonomous espionage, the traditional defense paradigms break down. We need to significantly expand our security and verification efforts to meet this moment.

Independent capability scaling tracking curves
Independent capability scaling tracking curves
UK AISI • Evaluation Intelligence

"The length of tasks frontier models can autonomously complete in our narrow cyber suite has been doubling every few months. This doubling rate has become faster over time, and recent models exceeded our previous trends."

Vulnerabilities & Attack Surfaces

Security challenges for frontier AI

Asset Protection

Confidentiality

A frontier model costs hundreds of millions of dollars to train, resulting in a few terabytes of weights. Unlike ordinary software secrets, leaked weights cannot be patched, recalled, or rotated.

Yet, those weights must be read constantly: sharded across training clusters, checkpointed to shared storage, and decrypted into GPU memory inside production environments optimized for latency.

The defender has to close every single route across clusters, storage, firmware, hypervisors, and user access layers. An attacker needs exactly one. The cost of failure is unusually final: a leaked model permanently hands frontier capability to whoever holds the file.

Pipeline Trust

Integrity

Keeping weights secure is only the baseline requirement. A resilient system also demands total integrity: absolute assurance that the model executing in production is the exact model that was trained and evaluated.

The pipeline stretches across thousands of machines and many hands. Without robust cryptographic integrity mechanisms, variants can be swapped, tampered with, or quietly fine-tuned.

A model whose weights never leak can still be compromised or degraded along the line, and the operator may be the absolute last to know.

Proof & Audit

Verifiability

The property that ties AI security together is verifiability: the ability for an external auditor, regulator, or counterparty to confirm what a model or cluster is actually doing, rather than taking the operator's word for it.

We must be able to verify exactly what was trained, on which physical chips, running where, and under what controls.

Today, these questions are answered almost entirely by trust. The reach of every policy and safety case relies on what can be proven rather than promised.

Join the mission