SciVeri-Bench asks whether AI agents can evaluate what constitutes good science as domain scientists do. We collect open-ended research tasks and the scientific judgments made as their outcomes evolve.
The evaluation target
Three scientific verification tasks.
Research agents produce papers and artifacts. Evaluator agents then make scientific judgments about those outcomes, with domain scientists providing the reference.
TASK 01
Scientist Critique Generation
Where does the science fall short?
Identify weaknesses in an AI-generated scientific outcome and support each criticism with evidence from the paper or research artifacts.
Given to the evaluator
An AI-generated paper and its research artifacts
Expected judgment
Evidence-backed weaknesses and concrete improvements
TASK 02
Scientist Rubric Generation
What did the rubric miss?
Recognize scientific problems that an existing rubric does not cover, then revise its criteria in light of the observed outcomes.
Given to the evaluator
An AI outcome, the current rubric, and scientific feedback
Expected judgment
A revised rubric with a rationale for the changes
TASK 03
Scientist Taste Evaluation
Which outcome is better science?
Compare scientific outcomes for novelty, meaning, and quality of evidence, and assess how an agent’s preferences align with scientists’ judgments.
Given to the evaluator
Two or three AI-generated scientific outcomes
Expected judgment
Pairwise preferences, including ties, with scientific reasoning
The benchmark studies agreement with scientists’ verification judgments. Successful execution or a high automated task reward alone does not establish the scientific quality of an outcome. Final evaluation metrics and results will accompany the benchmark release.
How the evidence is collected
The scientist-in-the-loop.
Scientists and research agents iteratively improve the investigation. Each round records the critiques, rubric revisions, and preferences needed to study scientific verification.
01
Propose
Domain scientist
An open scientific question, a runnable environment, and an initial rubric.
02
Investigate
Research agents
Experiments, a scientific paper, and the artifacts that support its claims.
03
Review
Domain scientist
Evidence-backed critiques and pairwise judgments of the scientific outcomes.
04
Refine
Scientist + agents
A revised rubric and scientific feedback inform the next research round.
The revised rubric, critiques, and previous outcomes become inputs to the next research round.
What scientists review
The papers and task-specific artifacts, their scientific weaknesses, and which outcomes offer more meaningful findings. Critiques point to supporting pages, figures, tables, or artifacts.
What the harness does
Run research agents in the prepared environment, collect their outcomes, and make the papers and feedback forms available through OpenReview. The next round receives the revised rubric and the scientific feedback.
Proposal review and iterative investigation establish a candidate task. Independent external review is a separate check on its scientific value.
PHASE 1
Task proposal & proposer-nominated review
The scientist submits a research question, its significance, required resources, and an initial rubric. Nominated reviewers from the same field help refine its novelty, open-endedness, and scientific meaning.
After proposal review, the task package is prepared for agent execution. Scientists review the resulting papers over multiple rounds and revise the rubric as new weaknesses emerge.
PHASE 2 · PLANNED
Independent external review
An external scientist in a related field, independent of the proposer and nominated reviewers, assesses the proposed task, its rubric, and the resulting findings.
The review considers whether the work is interesting, novel, and scientifically valuable enough for inclusion in SciVeri-Bench. The detailed external review procedure is still being developed.
What a proposal includes
The science comes first. A task manager can help translate the proposal into a runnable package.
Task name, objective, and importance
Why a scalar score is insufficient
Required data, software, and environment
An initial scientific evaluation rubric
Expected papers and research artifacts
Nominated reviewers with field expertise
Design principles
A benchmark grounded in research practice.
Open-ended science
A task’s outcomes cannot be adequately evaluated with a single scalar score. The scientific question leaves room for alternative approaches, discoveries, and expert interpretation.
Novel, frontier-relevant questions
Scientists propose problems with significance for their field. Review considers whether the direction is scientifically meaningful and has the potential for publication in a Nature-family journal.
Evaluation that learns from experiments
Rubrics begin as scientists’ best current criteria. Research outcomes reveal weaknesses that those criteria missed, prompting additions, revisions, and more informative experiments.
One interface for collaboration
OpenReview connects scientists and AI researchers. Scientists focus on questions and judgments; task managers and the harness handle execution and publish outcomes for review.
Have a scientific question to contribute?
Help define what good science means in your field.