Can AI agents
verify good science?
Can an AI agent tell whether another agent is doing good science? SciVeri-Bench evaluates scientific critiques, evolving rubrics, and judgments of research quality against those of domain scientists.
Read the project overviewOpen-ended science · Evolving rubrics · Scientist judgment
What SciVeri-Bench evaluates
Scientific outcomes
Papers, experiments & research artifacts
AI evaluator
Judgments to evaluate
Domain scientist
Reference judgments
Critique
Where does the science fall short?
Rubric
What did the rubric miss?
Taste
Which outcome is better science?
Can agents match scientists’ judgments?
On open-ended, novel scientific tasks.
- Verification tasks
- 3
- Task-proposal opt-ins
- 16
- Contributor institutions
- 11
- Countries represented
- 5
News / Upcoming presentation
SciVeri-Bench at Frontier Data Summit
October 8, 2026 · Scientist-Verified Benchmark
What we evaluate
Three ways to verify
good science.
Each AI outcome includes a paper and task-specific research artifacts. Evaluator agents critique those outcomes, revise the rubric, and compare the science. Scientists provide the reference judgments across multiple rounds.
Verification task 01
Where does the science fall short?
Identify weaknesses in an AI-generated scientific outcome and support each criticism with evidence from the paper or research artifacts.
- Input
- An AI-generated paper and its research artifacts
- Output
- Scientific weaknesses grounded in the paper and research artifacts
The verification task
Illustrative review example
Research outcome
A paper reports a promising material using a single simulation setting.
Critique response
The claim is broader than the evidence: test sensitivity to simulation conditions and compare against an established baseline.
An example of the task format; no benchmark result is implied.
Why verification?
Science evolves.
So should its evaluation.
For open-ended scientific questions, there is no known reference answer for a novel discovery. What counts as good science becomes clearer as experiments are conducted and results accumulate.
Read our design principles- 01
Find weaknesses in the results.
Scientists identify the gaps, missing evidence, and scientific weaknesses that experiments reveal.
- 02
Refine criteria as evidence accumulates.
New findings shape what good science means for a task. Scientists add, split, or edit rubric criteria as they learn.
- 03
Judge which findings make better science.
Scientists draw on domain expertise and scientific intuition to compare the novelty, evidence, and value of different outcomes.
One interface, with scientists in the loop
You bring the science.
We help run the experiments.
Propose and review tasks, inspect agent-generated papers, and refine what good science means—all through OpenReview. Task managers help prepare the code and operate the research agents.
From proposal review to task upload, validation, agent execution, and scientific feedback, the work stays in the same OpenReview thread.
The paper, reviews, and scores shown here illustrate the workflow.
Open-ended tasks
Does the outcome need scientific judgment?
Explain why a single scalar score cannot adequately capture the quality of the science.
Novel tasks
Would the result matter to your field?
Propose a question with the significance and research potential of work in a Nature-family journal.
Feasible tasks
What would it take to investigate?
Describe the data, tools, compute, and practical scope, from a tractable first experiment to a longer-term moonshot.
Built with scientists
Scientific judgment starts
with scientific expertise.
Researchers across disciplines have signed up to propose open-ended tasks rooted in their own fields.
Affiliations reported by task-proposal volunteers · Roster as of Oct 1, 2026.
Of scientists, by scientists, for scientists
Bring a question your field cares about.
Propose an open-ended research problem, help review the resulting science, and shape the criteria that the next generation of scientific agents will be judged by.