Skip to content
Scientist-Verified Benchmark

Can AI agents judge good science?

AI agents are doing research. Can they also verify if other agents are doing good science? SciVeri-Bench puts their scientific judgments to the test, with domain scientists as the reference.

Open-ended research · Evolving rubrics · Scientist judgment

The scientist-in-the-loop
  1. Propose

    01

    Domain scientist

    An open scientific question, a runnable environment, and an initial rubric.

  2. Investigate

    02

    Research agents

    Experiments, a scientific paper, and the artifacts that support its claims.

  3. Review

    03

    Domain scientist

    Evidence-backed critiques and pairwise judgments of the scientific outcomes.

  4. Refine

    04

    Scientist + agents

    A revised rubric and scientific feedback inform the next research round.

Scientific feedback shapes the next round
Verification tasks
3
Task-proposal opt-ins
16
Contributor institutions
11
Countries represented
5

What we evaluate

Three ways to verify
scientific work.

Each task asks an evaluator to make a judgment that scientists make in practice. We study whether agents can critique, refine evaluation criteria, and recognize valuable science.

Verification task 01

Where does the science fall short?

Identify weaknesses in an AI-generated scientific outcome and support each criticism with evidence from the paper or research artifacts.

Input
An AI-generated paper and its research artifacts
Output
Evidence-backed weaknesses and concrete improvements
Explore this verification task

Illustrative review example

Research outcome

A paper reports a promising material using a single simulation setting.

Critique response

The claim is broader than the evidence: test sensitivity to simulation conditions and compare against an established baseline.

An example of the task format; no benchmark result is implied.

Why verification?

Science evolves.
So should its evaluation.

A fixed score can miss what makes a discovery scientifically valuable. SciVeri-Bench studies verification as an ongoing part of research.

Read our design principles
  1. 01

    Scientific evidence reveals new weaknesses.

    Before an experiment runs, even an expert cannot anticipate every way a finding might fall short.

  2. 02

    A useful rubric evolves with the research.

    New outcomes expose gaps in existing criteria. Scientific feedback turns those gaps into better evaluation.

  3. 03

    Scientific value takes judgment.

    Novelty, evidence, and meaningful findings need the perspective of domain scientists, beyond a single scalar score.

Built with scientists

Scientific judgment starts
with scientific expertise.

Researchers across disciplines have signed up to propose open-ended tasks rooted in their own fields.

Meet the contributors
Brookhaven National Laboratory
Indian Institute of Technology Delhi
Institut de Biologie de l'École Normale Supérieure (IBENS)
Korea University
Massachusetts Institute of Technology
Shell
Technical University of Denmark
University of Colorado Boulder
University of Florida
University of Maryland, Baltimore County
University of Minnesota

Affiliations reported by task-proposal volunteers · Roster as of Oct 1, 2026.

Of scientists, by scientists, for scientists

Bring a question your field cares about.

Propose an open-ended research problem, help review the resulting science, and shape the criteria that the next generation of scientific agents will be judged by.

Become a contributor