News analysis medium confidence

METR time-horizon metric nears saturation; GDM's control roadmap proposes 80%-reliability measurement

Google DeepMind's AI Control Roadmap — the first detailed public control roadmap from a frontier lab — covers threat models for internally deployed AI, capability-level assessment, attack-execution evaluations, and mitigations (including in

On 2026-07-13, the verified AI news record added a significant safety, evals & benchmarks development: Google DeepMind's AI Control Roadmap — the first detailed public control roadmap from a frontier lab — covers threat models for internally deployed AI, capability-level assessment, attack-execution evaluations, and mitigations (including in-deployment honeypots once CoT monitoring is lost). It cites METR's 50%-reliability time-horizon suite as near-saturated (tasks up to ~16 human-hours; Claude Opus 4.6 ~14.5h) and proposes measuring at 80% reliability plus per-step hazard analysis; it concedes measuring attack-execution capability is an open problem and asks the external community for help.

Context

Google DeepMind's AI Control Roadmap — the first detailed public control roadmap from a frontier lab — covers threat models for internally deployed AI, capability-level assessment, attack-execution evaluations, and mitigations (including in-deployment honeypots once CoT monitoring is lost). It cites METR's 50%-reliability time-horizon suite as near-saturated (tasks up to ~16 human-hours; Claude Opus 4.6 ~14.5h) and proposes measuring at 80% reliability plus per-step hazard analysis; it concedes measuring attack-execution capability is an open problem and asks the external community for help. Doubling-time estimates now cited at ~4.2–7 months depending on source (americandefault.org/indicators/the-horizon) — treat as contested.

What changed

GDM proposes a control roadmap and flags attack-execution measurement as an open problem. According to arXiv 2607.13087v1 (PRIMARY, fetched), the supporting record states: “We have significant uncertainty about how to measure a model's attack execution capability, as this is an open problem in the field… We encourage the external research community to help advance the measurement science." And (D3 mitigations): "It may therefore make sense at this point to adopt in-deployment honeypots: fake opportunities for a misaligned model to sabotage the company or gain power.”.

Why it matters

Signals where lab safety cases and government pre-market eval requirements (EU, US framework) are heading; candid admission that key dangerous-capability metrics don't exist yet. The safety angle matters because evaluation quality, disclosure, and monitoring determine whether capability claims can be trusted.

Details

The research file records the item under “METR time-horizon metric nears saturation; GDM's control roadmap proposes 80%-reliability measurement” with source timing of GDM roadmap 2026-07-13 (arXiv 2607.13087). The captured research confidence note is: High (primary); Medium on specific saturation figures. Doubling-time estimates now cited at ~4.2–7 months depending on source (americandefault.org/indicators/the-horizon) — treat as contested.

Limitations and caveats

This item is based on a single authoritative source or a company-attributed claim captured in the research file; independent corroboration was not established in the research window.

Sources

Update note: Last reviewed 2026-07-22. Next checkpoint: monitor official channels and the linked source record.

Sources

Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.