News analysis high confidence

FLI Summer 2026 AI Safety Index: no lab above C+; red-line pledges weakened industry-wide

FLI's semiannual index graded 9 frontier developers on 37 indicators across 6 domains (panel: Russell, Krueger, Li, Maharaj, Revanur, Trager, Yi Zeng; evidence through June 3). Anthropic C+ (2.66, leads 5 of 6 domains), OpenAI C (2.28, lead

On 2026-07-07, the verified AI news record added a significant safety, evals & benchmarks development: FLI's semiannual index graded 9 frontier developers on 37 indicators across 6 domains (panel: Russell, Krueger, Li, Maharaj, Revanur, Trager, Yi Zeng; evidence through June 3). Anthropic C+ (2.66, leads 5 of 6 domains), OpenAI C (2.28, leads Risk Assessment), GDM C (2.01), Meta D+ (1.32, only riser), Z.ai D− (0.88), Alibaba Cloud D− (0.87), xAI F (0.65, 4th→7th), DeepSeek F (0.47), Mistral F (0.33, new, last). Findings: Anthropic/OpenAI/GDM/Meta weakened or voided unilateral pause pledges, replacing them with competitor-contingent conditions ("moving the goalposts"); Existential Safety weakest domain (no company above C−); reviewers flagged the industry's pivot to military AI; Mistral disputed the framework's fit for open models.

Context

FLI's semiannual index graded 9 frontier developers on 37 indicators across 6 domains (panel: Russell, Krueger, Li, Maharaj, Revanur, Trager, Yi Zeng; evidence through June 3). Anthropic C+ (2.66, leads 5 of 6 domains), OpenAI C (2.28, leads Risk Assessment), GDM C (2.01), Meta D+ (1.32, only riser), Z.ai D− (0.88), Alibaba Cloud D− (0.87), xAI F (0.65, 4th→7th), DeepSeek F (0.47), Mistral F (0.33, new, last). Findings: Anthropic/OpenAI/GDM/Meta weakened or voided unilateral pause pledges, replacing them with competitor-contingent conditions ("moving the goalposts"); Existential Safety weakest domain (no company above C−); reviewers flagged the industry's pivot to military AI; Mistral disputed the framework's fit for open models. Limitations: grades labs' policies/disclosures, not deployed products; Winter 2025→Summer 2026 trend table shows mostly downward drift (xAI −0.52, DeepSeek −0.55 sharpest). Note one secondary (theplanettools) misstates some scores (says "seven labs") — use the primary PDF/digitalapplied table.

What changed

Frontier labs retreated from unconditional pause commitments; inadequate safety is global. According to FLI AI Safety Index Summer 2026 2-pager PDF (PRIMARY, fetched); digitalapplied readout with full score table (2026-07-17), the supporting record states: “Even industry leaders in safety practices are retreating from prior commitments… Anthropic, OpenAI, Google DeepMind, and Meta have weakened or voided pledges to pause unilaterally if certain red lines were approached. Reviewers call this 'moving goalposts'… Existential Safety is the weakest domain industry-wide. No company exceeds C-; most score D or below.”.

Why it matters

The leading external lab-safety scorecard; the "conditional redlines" critique is shaping EU/US pre-market evaluation debates; seven of eight returning labs scored lower than in Winter 2025. The safety angle matters because evaluation quality, disclosure, and monitoring determine whether capability claims can be trusted.

Details

The research file records the item under “FLI Summer 2026 AI Safety Index: no lab above C+; red-line pledges weakened industry-wide” with source timing of 2026-07-07 (report; coverage Jul 7–17). The captured research confidence note is: High (primary). Additional captured source links are listed below so readers can inspect the evidence trail rather than rely on a single summary. Limitations: grades labs' policies/disclosures, not deployed products; Winter 2025→Summer 2026 trend table shows mostly downward drift (xAI −0.52, DeepSeek −0.55 sharpest). Note one secondary (theplanettools) misstates some scores (says "seven labs") — use the primary PDF/digitalapplied table.

Limitations and caveats

The research file did not identify a blocking caveat, but vendor-supplied claims should still be read as company statements unless independently confirmed.

Sources

Update note: Last reviewed 2026-07-22. Next checkpoint: monitor official channels and the linked source record.

Sources

Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.