Google DeepMind publishes first detailed AI control roadmap from a frontier lab
The arXiv paper lays out detection levels from chain-of-thought monitoring to opaque-reasoning regimes with honeypots — and concedes that measuring attack-execution capability remains an open problem.
Google DeepMind published its AI Control roadmap as an arXiv preprint on July 13, 2026 — the first detailed public control roadmap from a frontier AI lab. The paper sets out threat models for internally deployed AI systems, a framework for capability-level assessment, four detection levels (D1-D4) and corresponding response tiers.
Context
AI control — the discipline of maintaining safety even when models may be actively trying to subvert oversight — has moved from a niche research topic to a deployment concern as labs hand models access to internal infrastructure. The roadmap arrives amid a July cluster of safety-evaluation activity, including the Future of Life Institute's Summer 2026 AI Safety Index, and follows mounting evidence that frontier models can behave differently under evaluation.
What changed
Per the paper, the roadmap:
- Defines threat models for internally deployed AI, where models operate with meaningful access to lab systems.
- Introduces detection levels D1 through D4, spanning regimes where chain-of-thought monitoring remains viable up to opaque-reasoning regimes where mitigations such as in-deployment honeypots are required.
- Proposes capability-level assessment and response tiers tied to those detection levels.
- Cites METR's 50%-reliability time-horizon suite as near-saturated — with tasks up to roughly 16 human-hours and Claude Opus 4.6 at about 14.5 hours — and proposes measuring at 80% reliability instead, plus per-step hazard analysis.
The paper is candid about its limits: "We have significant uncertainty about how to measure a model's attack execution capability, as this is an open problem in the field… We encourage the external research community to help advance the measurement science."
Why it matters
Control has largely been discussed through conceptual papers and small-scale demonstrations; a frontier lab publishing its internal roadmap — with named detection regimes, explicit thresholds and admitted blind spots — gives regulators, evaluators and competing labs a concrete reference architecture. The proposal to shift time-horizon measurement from 50% to 80% reliability is a substantive methodological challenge to the field's dominant capability metric at the moment that metric is nearing saturation.
Details
The roadmap treats attack-execution capability — whether a model could actually carry out a harmful plan against lab defenses — as distinct from the task-completion capabilities measured by standard benchmarks, and states plainly that the field lacks validated measurement science for it. The METR time-horizon numbers the paper cites remain contested on methodology grounds, which the shift to 80% reliability is partly designed to address. One secondary source dated a GDM control framework to June 18; the arXiv v1 is July 13, which this report uses.
Limitations and caveats
The roadmap is a framework document, not evidence that mitigations work at the described detection levels; its authors flag attack-execution measurement as unsolved. The METR figures cited within are contested and are quoted here as the paper presents them. This report is based on the arXiv HTML v1 as the primary source; subsequent revisions may change details.
Sources
*Update note: This post was last reviewed on 2026-07-22. Watch for revised versions of the preprint and for external evaluations responding to the paper's call for measurement work.*
Sources
- arXiv 2607.13087v1 — research
Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.