Skip to content
Deploy-first · Scientist-driven

Can AI agents verify good science?

Can an AI agent tell whether another agent is doing good science? SciVeri-Bench evaluates scientific critiques, evolving rubrics, and judgments of research quality against those of domain scientists.

Read the project overview

Open-ended science · Evolving rubrics · Scientist judgment

What SciVeri-Bench evaluates

Scientific outcomes

Papers, experiments & research artifacts

AI evaluator

Judgments to evaluate

Domain scientist

Reference judgments

Can agents match scientists’ judgments?

On open-ended, novel scientific tasks.

Verification tasks
3
Task-proposal opt-ins
16
Contributor institutions
11
Countries represented
5

News / Upcoming presentation

SciVeri-Bench at Frontier Data Summit

October 8, 2026 · Scientist-Verified Benchmark

Read the news

What we evaluate

Three ways to verify
good science.

Each AI outcome includes a paper and task-specific research artifacts. Evaluator agents critique those outcomes, revise the rubric, and compare the science. Scientists provide the reference judgments across multiple rounds.

Verification task 01

Where does the science fall short?

Identify weaknesses in an AI-generated scientific outcome and support each criticism with evidence from the paper or research artifacts.

Input
An AI-generated paper and its research artifacts
Output
Scientific weaknesses grounded in the paper and research artifacts
Explore this verification task

The verification task

AI outcome
AI evaluator
Weaknesses
Compared with the scientist’s critique

Illustrative review example

Research outcome

A paper reports a promising material using a single simulation setting.

Critique response

The claim is broader than the evidence: test sensitivity to simulation conditions and compare against an established baseline.

An example of the task format; no benchmark result is implied.

Why verification?

Science evolves.
So should its evaluation.

For open-ended scientific questions, there is no known reference answer for a novel discovery. What counts as good science becomes clearer as experiments are conducted and results accumulate.

Read our design principles
  1. 01

    Find weaknesses in the results.

    Scientists identify the gaps, missing evidence, and scientific weaknesses that experiments reveal.

  2. 02

    Refine criteria as evidence accumulates.

    New findings shape what good science means for a task. Scientists add, split, or edit rubric criteria as they learn.

  3. 03

    Judge which findings make better science.

    Scientists draw on domain expertise and scientific intuition to compare the novelty, evidence, and value of different outcomes.

One interface, with scientists in the loop

You bring the science.
We help run the experiments.

Propose and review tasks, inspect agent-generated papers, and refine what good science means—all through OpenReview. Task managers help prepare the code and operate the research agents.

From proposal review to task upload, validation, agent execution, and scientific feedback, the work stays in the same OpenReview thread.

The paper, reviews, and scores shown here illustrate the workflow.

Open-ended tasks

Does the outcome need scientific judgment?

Explain why a single scalar score cannot adequately capture the quality of the science.

Novel tasks

Would the result matter to your field?

Propose a question with the significance and research potential of work in a Nature-family journal.

Feasible tasks

What would it take to investigate?

Describe the data, tools, compute, and practical scope, from a tractable first experiment to a longer-term moonshot.

Built with scientists

Scientific judgment starts
with scientific expertise.

Researchers across disciplines have signed up to propose open-ended tasks rooted in their own fields.

Meet the contributors
Brookhaven National Laboratory
Indian Institute of Technology Delhi
Institut de Biologie de l'École Normale Supérieure (IBENS)
Korea University
Massachusetts Institute of Technology
Shell
Technical University of Denmark
University of Colorado Boulder
University of Florida
University of Maryland, Baltimore County
University of Minnesota

Affiliations reported by task-proposal volunteers · Roster as of Oct 1, 2026.

Of scientists, by scientists, for scientists

Bring a question your field cares about.

Propose an open-ended research problem, help review the resulting science, and shape the criteria that the next generation of scientific agents will be judged by.

Become a contributor