News confirmed high confidence

GeneBench-Pro Tests Whether AI Can Make the Messy Decisions Biology Requires

OpenAI's 129-problem benchmark asks agents to choose and revise real analysis workflows—not simply recall scientific facts—and the leading reported result still solves fewer than one in three problems.

Scientific data rarely arrive with a neat question and a clean spreadsheet. A researcher has to decide what the data can support, notice when an assumption has failed, change methods, and know when a result is strong enough to guide the next decision.

GeneBench-Pro is OpenAI's attempt to test that difficult middle of the scientific process. Released on June 30 with a bioRxiv preprint, the benchmark gives AI agents 129 computational-biology problems spanning 10 domains and 21 subdomains, including statistical genetics, cancer genomics, proteomics and clinical diagnostics.

What the benchmark measures

Each problem places an agent in an isolated workspace with a short prompt, data files and a standard bioinformatics toolset. The agent must inspect the data, select an analytical path, respond to diagnostic evidence and return a structured conclusion.

The benchmark is not built from ordinary historical datasets. OpenAI says each problem is created synthetically, with a known causal structure and a directly simulated data-generating process. That lets the authors grade answers against known targets and test whether a superficially plausible but analytically wrong approach fails.

This design addresses a persistent benchmark problem: two researchers can make different but defensible choices on real, messy data. If both choices are reasonable, a single “correct” answer may measure the benchmark author's preference rather than scientific judgment.

What was released

OpenAI has published 10 representative problems with an interactive browser and a Hugging Face package. External domain experts reviewed 82 of the 129 questions. The company also says it will provide a 50-question subset to Artificial Analysis for independent third-party benchmarking; those independent results were not part of the initial release.

OpenAI reports that GPT-5.6 Sol passes 28.7% of the full suite at its highest reasoning level, rising to 31.5% in separately reported Pro runs. Those are vendor-reported results from a benchmark created by OpenAI, so they should not be read as an independent model ranking.

The low absolute result is nevertheless important. Even the strongest reported system fails more than two-thirds of the problems. The gap is not simply missing biological knowledge: the paper says agents often notice a local warning sign but fail to carry its consequences through the rest of the analysis.

Why it matters

Many AI-for-science claims focus on speed—how quickly a model can write code, search papers or run a familiar analysis. GeneBench-Pro asks a harder question: can it exercise enough judgment to choose the right analysis in the first place?

Synthetic construction gives the benchmark a defensible answer key, but it also creates a limitation. Real laboratories contain missing records, institutional habits and experimental surprises that no simulation captures perfectly. Ten public problems are enough to inspect the format, not enough to reproduce the full result.

That makes the forthcoming third-party track the next meaningful checkpoint. If outside evaluators can reproduce the scoring and models improve without learning the hidden suite, GeneBench-Pro could become a useful measure of scientific agents. Until then, it is a promising benchmark and a vendor-authored claim about the systems it tests.

Sources

Update note: Published and source-checked on 2026-07-24. Next checkpoint: independent results from the 50-question Artificial Analysis subset.

Sources

Drafted with AI assistance from verified source material and reviewed for factual accuracy, attribution, clarity, and label integrity.