WC2026-Agents: a contamination-free forecasting benchmark built on the 2026 FIFA World Cup
Four frontier agents (Claude Opus 4.8, GPT-5.5 high, Gemini 3.1 Pro, Grok Expert Mode) ran an identical search–act–reflect loop with virtual $100 bets on all 104 World Cup matches (all post-training-cutoff, hence contamination-free by const
On 2026-07-20, the verified AI news record added a significant safety, evals & benchmarks development: Four frontier agents (Claude Opus 4.8, GPT-5.5 high, Gemini 3.1 Pro, Grok Expert Mode) ran an identical search–act–reflect loop with virtual $100 bets on all 104 World Cup matches (all post-training-cutoff, hence contamination-free by construction), benchmarked against betting-market odds. Findings: identical top pick in 92% of matches; no agent beats the market's Brier score (market 0.469 vs best agent 0.471); betting ROI −18% to +10%; self-reported error rates on wrong picks 36%–86%. Release: 416 forecasts, 414 reflections with verbatim reasoning, odds, code (CC BY 4.0 / MIT).
Context
Four frontier agents (Claude Opus 4.8, GPT-5.5 high, Gemini 3.1 Pro, Grok Expert Mode) ran an identical search–act–reflect loop with virtual $100 bets on all 104 World Cup matches (all post-training-cutoff, hence contamination-free by construction), benchmarked against betting-market odds. Findings: identical top pick in 92% of matches; no agent beats the market's Brier score (market 0.469 vs best agent 0.471); betting ROI −18% to +10%; self-reported error rates on wrong picks 36%–86%. Release: 416 forecasts, 414 reflections with verbatim reasoning, odds, code (CC BY 4.0 / MIT). Authors' own limitations: single tournament, wide intervals; consumer interfaces (models may update mid-window); odds from mostly one book. Counters "wisdom of the silicon crowd": shared retrieval surfaces mean shared errors.
What changed
Contamination-free design; agents converge on market prices and none beats the market. According to arXiv 2607.17765 (PRIMARY, fetched), the supporting record states: “the four agents issue an identical top pick in 92% of matches and none beats the market's Brier score, yet they diverge sharply as decision-makers… The benchmark thus measures calibration, decision quality, and self-knowledge — axes on which frontier models differ even when their predictions do not.”.
Why it matters
A working template for live, rolling, contamination-proof agent evaluation that measures calibration, decision quality, and self-knowledge instead of static accuracy — the field's direct answer to the contamination crisis in items 7–8. The safety angle matters because evaluation quality, disclosure, and monitoring determine whether capability claims can be trusted.
Details
The research file records the item under “WC2026-Agents: a contamination-free forecasting benchmark built on the 2026 FIFA World Cup” with source timing of arXiv v1 2026-07-20 (2607.17765); KDD '27 Datasets & Benchmarks track. The captured research confidence note is: High (primary). Additional captured source links are listed below so readers can inspect the evidence trail rather than rely on a single summary. Authors' own limitations: single tournament, wide intervals; consumer interfaces (models may update mid-window); odds from mostly one book. Counters "wisdom of the silicon crowd": shared retrieval surfaces mean shared errors.
Limitations and caveats
The research file did not identify a blocking caveat, but vendor-supplied claims should still be read as company statements unless independently confirmed.
Sources
Update note: Last reviewed 2026-07-22. Next checkpoint: monitor official channels and the linked source record.
Sources
- arXiv 2607.17765 (PRIMARY, fetched) — research
- arXiv 2607.17765 (PRIMARY, fetched) — official
Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.