Can AI agents
judge good science?
AI agents are doing research. Can they also verify if other agents are doing good science? SciVeri-Bench puts their scientific judgments to the test, with domain scientists as the reference.
Open-ended research · Evolving rubrics · Scientist judgment
Propose
01Domain scientist
An open scientific question, a runnable environment, and an initial rubric.
Investigate
02Research agents
Experiments, a scientific paper, and the artifacts that support its claims.
Review
03Domain scientist
Evidence-backed critiques and pairwise judgments of the scientific outcomes.
Refine
04Scientist + agents
A revised rubric and scientific feedback inform the next research round.
- Verification tasks
- 3
- Task-proposal opt-ins
- 16
- Contributor institutions
- 11
- Countries represented
- 5
What we evaluate
Three ways to verify
scientific work.
Each task asks an evaluator to make a judgment that scientists make in practice. We study whether agents can critique, refine evaluation criteria, and recognize valuable science.
Verification task 01
Where does the science fall short?
Identify weaknesses in an AI-generated scientific outcome and support each criticism with evidence from the paper or research artifacts.
- Input
- An AI-generated paper and its research artifacts
- Output
- Evidence-backed weaknesses and concrete improvements
Illustrative review example
Research outcome
A paper reports a promising material using a single simulation setting.
Critique response
The claim is broader than the evidence: test sensitivity to simulation conditions and compare against an established baseline.
An example of the task format; no benchmark result is implied.
Why verification?
Science evolves.
So should its evaluation.
A fixed score can miss what makes a discovery scientifically valuable. SciVeri-Bench studies verification as an ongoing part of research.
Read our design principles- 01
Scientific evidence reveals new weaknesses.
Before an experiment runs, even an expert cannot anticipate every way a finding might fall short.
- 02
A useful rubric evolves with the research.
New outcomes expose gaps in existing criteria. Scientific feedback turns those gaps into better evaluation.
- 03
Scientific value takes judgment.
Novelty, evidence, and meaningful findings need the perspective of domain scientists, beyond a single scalar score.
Built with scientists
Scientific judgment starts
with scientific expertise.
Researchers across disciplines have signed up to propose open-ended tasks rooted in their own fields.
Affiliations reported by task-proposal volunteers · Roster as of Oct 1, 2026.
Of scientists, by scientists, for scientists
Bring a question your field cares about.
Propose an open-ended research problem, help review the resulting science, and shape the criteria that the next generation of scientific agents will be judged by.