test-lab.ai mid-2026 browser-agent benchmark: routine plans saturated, field splits on hard journeys
test-lab.ai's mid-year sweep added Claude Fable 5, Opus 4.8, GPT-5.5 to its production browser-automation harness. On the routine plan, five models all pass 100% ("the 2026 frontier buys you nothing"; 60x cost spread for the same green chec
On 2026-07-19, the verified AI news record added a significant safety, evals & benchmarks development: test-lab.ai's mid-year sweep added Claude Fable 5, Opus 4.8, GPT-5.5 to its production browser-automation harness. On the routine plan, five models all pass 100% ("the 2026 frontier buys you nothing"; 60x cost spread for the same green checkmark). On a new 30+ step hard plan: Fable 5 92%, Opus 4.8 74%, GPT-5.5 66%, Gemini 3.1 Pro 58%. GPT-5.6 missed the cutoff and joins the next sweep. Same-harness, serial, no-prompt-tuning methodology.
Context
test-lab.ai's mid-year sweep added Claude Fable 5, Opus 4.8, GPT-5.5 to its production browser-automation harness. On the routine plan, five models all pass 100% ("the 2026 frontier buys you nothing"; 60x cost spread for the same green checkmark). On a new 30+ step hard plan: Fable 5 92%, Opus 4.8 74%, GPT-5.5 66%, Gemini 3.1 Pro 58%. GPT-5.6 missed the cutoff and joins the next sweep. Same-harness, serial, no-prompt-tuning methodology. Caveat: vendor that sells model-routing; single target surface; GPT-5.6 absent.
What changed
Routine browser tasks no longer separate frontier models; hard journeys do. According to test-lab.ai (vendor-run independent harness; Authority NA but methodology transparent), the supporting record states: “on a routine plan, the 2026 frontier buys you nothing. Five models pass 100% of runs. The cheapest of them costs 3 cents a run and the most expensive costs nearly four dollars, a 60x spread for the same green checkmark.”.
Why it matters
One of the few serial independent evals tracking real agentic reliability; documents benchmark saturation in production-adjacent tasks and the "route by difficulty" deployment answer. The safety angle matters because evaluation quality, disclosure, and monitoring determine whether capability claims can be trusted.
Details
The research file records the item under “test-lab.ai mid-2026 browser-agent benchmark: routine plans saturated, field splits on hard journeys” with source timing of 2026-07-19. The captured research confidence note is: Medium. Caveat: vendor that sells model-routing; single target surface; GPT-5.6 absent.
Limitations and caveats
This item is based on a single authoritative source or a company-attributed claim captured in the research file; independent corroboration was not established in the research window.
Sources
Update note: Last reviewed 2026-07-22. Next checkpoint: monitor official channels and the linked source record.
Sources
Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.