News confirmed high confidence

Google DeepMind publishes first detailed AI control roadmap from a frontier lab

The arXiv paper lays out detection levels from chain-of-thought monitoring to opaque-reasoning regimes with honeypots — and concedes that measuring attack-execution capability remains an open problem.

Google DeepMind published its AI Control roadmap as an arXiv preprint on July 13, 2026 — the first detailed public control roadmap from a frontier AI lab. The paper sets out threat models for internally deployed AI systems, a framework for capability-level assessment, four detection levels (D1-D4) and corresponding response tiers.

Context

AI control — the discipline of maintaining safety even when models may be actively trying to subvert oversight — has moved from a niche research topic to a deployment concern as labs hand models access to internal infrastructure. The roadmap arrives amid a July cluster of safety-evaluation activity, including the Future of Life Institute's Summer 2026 AI Safety Index, and follows mounting evidence that frontier models can behave differently under evaluation.

What changed

Per the paper, the roadmap:

The paper is candid about its limits: "We have significant uncertainty about how to measure a model's attack execution capability, as this is an open problem in the field… We encourage the external research community to help advance the measurement science."

Why it matters

Control has largely been discussed through conceptual papers and small-scale demonstrations; a frontier lab publishing its internal roadmap — with named detection regimes, explicit thresholds and admitted blind spots — gives regulators, evaluators and competing labs a concrete reference architecture. The proposal to shift time-horizon measurement from 50% to 80% reliability is a substantive methodological challenge to the field's dominant capability metric at the moment that metric is nearing saturation.

Details

The roadmap treats attack-execution capability — whether a model could actually carry out a harmful plan against lab defenses — as distinct from the task-completion capabilities measured by standard benchmarks, and states plainly that the field lacks validated measurement science for it. The METR time-horizon numbers the paper cites remain contested on methodology grounds, which the shift to 80% reliability is partly designed to address. One secondary source dated a GDM control framework to June 18; the arXiv v1 is July 13, which this report uses.

Limitations and caveats

The roadmap is a framework document, not evidence that mitigations work at the described detection levels; its authors flag attack-execution measurement as unsolved. The METR figures cited within are contested and are quoted here as the paper presents them. This report is based on the arXiv HTML v1 as the primary source; subsequent revisions may change details.

Sources

*Update note: This post was last reviewed on 2026-07-22. Watch for revised versions of the preprint and for external evaluations responding to the paper's call for measurement work.*

Sources

Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.