Scientist Critique Generation
Identify weaknesses in an AI-generated scientific outcome and support each criticism with evidence from the paper or research artifacts.
Results pending
Scientific verification results
The SciVeri-Bench evaluation will compare agents’ scientific judgments with those of domain scientists across three verification tasks.
Task collection and scientific review are in progress. The evaluation protocol, reference judgments, and measured results will be reported with the benchmark release.
Identify weaknesses in an AI-generated scientific outcome and support each criticism with evidence from the paper or research artifacts.
Results pending
Recognize scientific problems that an existing rubric does not cover, then revise its criteria in light of the observed outcomes.
Results pending
Compare scientific outcomes for novelty, meaning, and quality of evidence, and assess how an agent’s preferences align with scientists’ judgments.
Results pending
Supporting exploration
The earlier workflow interface is retained below as a synthetic demonstration. Its generated scores and performance trends are separate from SciVeri-Bench’s scientific verification evaluation.
Illustrative, generated data. These tables and charts contain no measured SciVeri-Bench verification results.
Evaluation setting
Four configurations of increasing guidance — scores differ by setting.
7 agents · ranked by Rubric Score. Click a metric header to re-sort.
| # | Agent Harness | Model | Tasks | ||
|---|---|---|---|---|---|
Claude Code | Claude Opus 4.8 Anthropic | 17 | 77.1 | 100.0 | |
Codex | GPT-5.5 OpenAI | 17 | 74.8 | 92.0 | |
OpenHands | Claude Opus 4.8 All Hands AI | 17 | 74.5 | 93.9 | |
| 4 | Gemini CLI | Gemini 3.1 Pro Google | 17 | 71.5 | 90.8 |
| 5 | OpenHands | GPT-5.5 All Hands AI | 17 | 69.6 | 89.4 |
| 6 | OpenHands | Gemini 3.1 Pro All Hands AI | 17 | 68.9 | 87.5 |
| 7 | OpenHands | Qwen3.5-397B-A17B Alibaba | 17 | 63.8 | 82.1 |
Evolutionary loop
Illustrative projectionA predicted trajectory for an agent solving Gangmin Son’s proposed spin-glass task in an evolutionary loop: each iteration proposes a candidate pipeline, and the best-so-far Outcome Verification score only ever steps up. Every jump is a methodological breakthrough — hover a point, or pick a jump below, to see what changed. This is a hypothetical projection, not measured data.
Running-max envelope (blue) over evolutionary candidates (grey)
The extracted upper critical dimension and its confidence interval land consistent with the loop-expansion prediction D_U ≤ 8 [Angelini et al., 2022] — above the classical D_U = 6. The residual gap to a perfect score is the open-problem uncertainty: there is no ground-truth D_U to score against.