Learn learning medium confidence

[LEARNING/RESOURCE] Terminal-Bench 2.1 explainer + live leaderboard (benchmark literacy)

Two practitioner references frame July's headline agentic metric: qaskills.sh's guide explains Terminal-Bench's end-state verification design (Docker sandbox, pytest-style checks of machine final state, not transcripts; Stanford × Laude Ins

On 2026-07-21, the verified AI news record added a significant learning, rumors & expectations development: Two practitioner references frame July's headline agentic metric: qaskills.sh's guide explains Terminal-Bench's end-state verification design (Docker sandbox, pytest-style checks of machine final state, not transcripts; Stanford × Laude Institute), and codingfleet's leaderboard (updated Jul 21) tracks 2.1 scores — GPT-5.6 Sol 88.8% single-agent, Kimi K3 88.3% (vendor table), Muse Spark 1.1 80.0%, Gemini 3.6 Flash 78.0% — with an explicit warning that 2.1 and 2.0 scores are NOT comparable and vendor/leaderboard harnesses differ.

Context

Two practitioner references frame July's headline agentic metric: qaskills.sh's guide explains Terminal-Bench's end-state verification design (Docker sandbox, pytest-style checks of machine final state, not transcripts; Stanford × Laude Institute), and codingfleet's leaderboard (updated Jul 21) tracks 2.1 scores — GPT-5.6 Sol 88.8% single-agent, Kimi K3 88.3% (vendor table), Muse Spark 1.1 80.0%, Gemini 3.6 Flash 78.0% — with an explicit warning that 2.1 and 2.0 scores are NOT comparable and vendor/leaderboard harnesses differ. Directly supports editorial benchmark-literacy coverage of the GPT-5.6 Terminal-Bench record claims (wide01 item 1) and Kimi K3's launch-table numbers. Model for "how to read a benchmark" explainers.

What changed

Two practitioner references frame July's headline agentic metric: qaskills.sh's guide explains Terminal-Bench's end-state verification design (Docker sandbox, pytest-style checks of machine final state, not transcripts; Stanford × Laude Institute), and codingfleet's leaderboard (updated Jul 21) tracks 2.1 scores — GPT-5.6 Sol 88.8% single-agent, Kimi K3 88.3% (vendor table), Muse Spark 1.1 80.0%, Gemini 3.6 Flash 78.0% — with an explicit warning that 2.1 and 2.0 scores are NOT comparable and vendor/leaderboard harnesses differ. According to qaskills.sh (trade explainer); codingfleet.com (leaderboard aggregator), the supporting record states: “Scores NOT directly comparable across versions — 2.1 is harder… Results combine vendor and leaderboard harnesses. Compare directionally unless the same evaluator and scaffold were used.”.

Why it matters

Directly supports editorial benchmark-literacy coverage of the GPT-5.6 Terminal-Bench record claims (wide01 item 1) and Kimi K3's launch-table numbers. Model for "how to read a benchmark" explainers. The signal matters only if readers can separate verified evidence from expectation, which is why the sourcing and next checkpoint are explicit.

Details

The research file records the item under “[LEARNING/RESOURCE] Terminal-Bench 2.1 explainer + live leaderboard (benchmark literacy)” with source timing of Guide 2026-06-15; leaderboard updated 2026-07-21. The captured research confidence note is: High (resource exists); Medium on individual vendor-reported scores | Status: learning/resource. Additional captured source links are listed below so readers can inspect the evidence trail rather than rely on a single summary. Directly supports editorial benchmark-literacy coverage of the GPT-5.6 Terminal-Bench record claims (wide01 item 1) and Kimi K3's launch-table numbers. Model for "how to read a benchmark" explainers.

Limitations and caveats

This item is based on a single authoritative source or a company-attributed claim captured in the research file; independent corroboration was not established in the research window.

Sources

Update note: Last reviewed 2026-07-22. Next checkpoint: monitor official channels and the linked source record.

Sources

Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.