News confirmed medium confidence

test-lab.ai mid-2026 browser-agent benchmark: routine plans saturated, field splits on hard journeys

test-lab.ai's mid-year sweep added Claude Fable 5, Opus 4.8, GPT-5.5 to its production browser-automation harness. On the routine plan, five models all pass 100% ("the 2026 frontier buys you nothing"; 60x cost spread for the same green chec

On 2026-07-19, the verified AI news record added a significant safety, evals & benchmarks development: test-lab.ai's mid-year sweep added Claude Fable 5, Opus 4.8, GPT-5.5 to its production browser-automation harness. On the routine plan, five models all pass 100% ("the 2026 frontier buys you nothing"; 60x cost spread for the same green checkmark). On a new 30+ step hard plan: Fable 5 92%, Opus 4.8 74%, GPT-5.5 66%, Gemini 3.1 Pro 58%. GPT-5.6 missed the cutoff and joins the next sweep. Same-harness, serial, no-prompt-tuning methodology.

Context

test-lab.ai's mid-year sweep added Claude Fable 5, Opus 4.8, GPT-5.5 to its production browser-automation harness. On the routine plan, five models all pass 100% ("the 2026 frontier buys you nothing"; 60x cost spread for the same green checkmark). On a new 30+ step hard plan: Fable 5 92%, Opus 4.8 74%, GPT-5.5 66%, Gemini 3.1 Pro 58%. GPT-5.6 missed the cutoff and joins the next sweep. Same-harness, serial, no-prompt-tuning methodology. Caveat: vendor that sells model-routing; single target surface; GPT-5.6 absent.

What changed

Routine browser tasks no longer separate frontier models; hard journeys do. According to test-lab.ai (vendor-run independent harness; Authority NA but methodology transparent), the supporting record states: “on a routine plan, the 2026 frontier buys you nothing. Five models pass 100% of runs. The cheapest of them costs 3 cents a run and the most expensive costs nearly four dollars, a 60x spread for the same green checkmark.”.

Why it matters

One of the few serial independent evals tracking real agentic reliability; documents benchmark saturation in production-adjacent tasks and the "route by difficulty" deployment answer. The safety angle matters because evaluation quality, disclosure, and monitoring determine whether capability claims can be trusted.

Details

The research file records the item under “test-lab.ai mid-2026 browser-agent benchmark: routine plans saturated, field splits on hard journeys” with source timing of 2026-07-19. The captured research confidence note is: Medium. Caveat: vendor that sells model-routing; single target surface; GPT-5.6 absent.

Limitations and caveats

This item is based on a single authoritative source or a company-attributed claim captured in the research file; independent corroboration was not established in the research window.

Sources

Update note: Last reviewed 2026-07-22. Next checkpoint: monitor official channels and the linked source record.

Sources

Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.