Skip to content

Scientific verification results

How well do agents
judge the science?

The SciVeri-Bench evaluation will compare agents’ scientific judgments with those of domain scientists across three verification tasks.

Benchmark results are not yet published.

Task collection and scientific review are in progress. The evaluation protocol, reference judgments, and measured results will be reported with the benchmark release.

Evaluation design
01

Scientist Critique Generation

Identify weaknesses in an AI-generated scientific outcome and support each criticism with evidence from the paper or research artifacts.

Results pending

02

Scientist Rubric Generation

Recognize scientific problems that an existing rubric does not cover, then revise its criteria in light of the observed outcomes.

Results pending

03

Scientist Taste Evaluation

Compare scientific outcomes for novelty, meaning, and quality of evidence, and assess how an agent’s preferences align with scientists’ judgments.

Results pending

Supporting exploration

Workflow score demonstration

The earlier workflow interface is retained below as a synthetic demonstration. Its generated scores and performance trends are separate from SciVeri-Bench’s scientific verification evaluation.

View the legacy workflow demo · Synthetic data

Illustrative, generated data. These tables and charts contain no measured SciVeri-Bench verification results.

Evaluation setting

Four configurations of increasing guidance — scores differ by setting.

Overall ranking

7 agents · ranked by Rubric Score. Click a metric header to re-sort.

≥ 100 = threshold met
#Agent HarnessModel Tasks
Claude Code
Claude Opus 4.8
Anthropic
17
77.1
100.0
Codex
GPT-5.5
OpenAI
17
74.8
92.0
OpenHands
Claude Opus 4.8
All Hands AI
17
74.5
93.9
4
Gemini CLI
Gemini 3.1 Pro
Google
17
71.5
90.8
5
OpenHands
GPT-5.5
All Hands AI
17
69.6
89.4
6
OpenHands
Gemini 3.1 Pro
All Hands AI
17
68.9
87.5
7
OpenHands
Qwen3.5-397B-A17B
Alibaba
17
63.8
82.1

Evolutionary loop

Illustrative projection

Climbing a frontier task, jump by jump

A predicted trajectory for an agent solving Gangmin Son’s proposed spin-glass task in an evolutionary loop: each iteration proposes a candidate pipeline, and the best-so-far Outcome Verification score only ever steps up. Every jump is a methodological breakthrough — hover a point, or pick a jump below, to see what changed. This is a hypothetical projection, not measured data.

Best-so-far Outcome Verification score

Running-max envelope (blue) over evolutionary candidates (grey)

Jump 6 / 6iteration 21

Estimate consistent with D_U ≤ 8

Analytic-bound consistency+6% TCS

The extracted upper critical dimension and its confidence interval land consistent with the loop-expansion prediction D_U ≤ 8 [Angelini et al., 2022] — above the classical D_U = 6. The residual gap to a perfect score is the open-problem uncertainty: there is no ground-truth D_U to score against.

Best-so-far84%