Skip to content

The Scientist-Verified Benchmark

Can verification itself
be verified?

SciVeri-Bench asks whether AI agents can evaluate what constitutes good science as domain scientists do. We collect open-ended research tasks and the scientific judgments made as their outcomes evolve.

The evaluation target

Three scientific verification tasks.

Research agents produce papers and artifacts. Evaluator agents then make scientific judgments about those outcomes, with domain scientists providing the reference.

TASK 01

Scientist Critique Generation

Where does the science fall short?

Identify weaknesses in an AI-generated scientific outcome and support each criticism with evidence from the paper or research artifacts.

Given to the evaluator
An AI-generated paper and its research artifacts
Expected judgment
Evidence-backed weaknesses and concrete improvements

TASK 02

Scientist Rubric Generation

What did the rubric miss?

Recognize scientific problems that an existing rubric does not cover, then revise its criteria in light of the observed outcomes.

Given to the evaluator
An AI outcome, the current rubric, and scientific feedback
Expected judgment
A revised rubric with a rationale for the changes

TASK 03

Scientist Taste Evaluation

Which outcome is better science?

Compare scientific outcomes for novelty, meaning, and quality of evidence, and assess how an agent’s preferences align with scientists’ judgments.

Given to the evaluator
Two or three AI-generated scientific outcomes
Expected judgment
Pairwise preferences, including ties, with scientific reasoning

The benchmark studies agreement with scientists’ verification judgments. Successful execution or a high automated task reward alone does not establish the scientific quality of an outcome. Final evaluation metrics and results will accompany the benchmark release.

How the evidence is collected

The scientist-in-the-loop.

Scientists and research agents iteratively improve the investigation. Each round records the critiques, rubric revisions, and preferences needed to study scientific verification.

  1. 01

    Propose

    Domain scientist

    An open scientific question, a runnable environment, and an initial rubric.

  2. 02

    Investigate

    Research agents

    Experiments, a scientific paper, and the artifacts that support its claims.

  3. 03

    Review

    Domain scientist

    Evidence-backed critiques and pairwise judgments of the scientific outcomes.

  4. 04

    Refine

    Scientist + agents

    A revised rubric and scientific feedback inform the next research round.

The revised rubric, critiques, and previous outcomes become inputs to the next research round.

What scientists review

The papers and task-specific artifacts, their scientific weaknesses, and which outcomes offer more meaningful findings. Critiques point to supporting pages, figures, tables, or artifacts.

What the harness does

Run research agents in the prepared environment, collect their outcomes, and make the papers and feedback forms available through OpenReview. The next round receives the revised rubric and the scientific feedback.

Task collection procedure

Two phases of scientific review.

Proposal review and iterative investigation establish a candidate task. Independent external review is a separate check on its scientific value.

PHASE 1

Task proposal &
proposer-nominated review

The scientist submits a research question, its significance, required resources, and an initial rubric. Nominated reviewers from the same field help refine its novelty, open-endedness, and scientific meaning.

After proposal review, the task package is prepared for agent execution. Scientists review the resulting papers over multiple rounds and revise the rubric as new weaknesses emerge.

PHASE 2 · PLANNED

Independent
external review

An external scientist in a related field, independent of the proposer and nominated reviewers, assesses the proposed task, its rubric, and the resulting findings.

The review considers whether the work is interesting, novel, and scientifically valuable enough for inclusion in SciVeri-Bench. The detailed external review procedure is still being developed.

What a proposal includes

The science comes first. A task manager can help translate the proposal into a runnable package.

  • Task name, objective, and importance
  • Why a scalar score is insufficient
  • Required data, software, and environment
  • An initial scientific evaluation rubric
  • Expected papers and research artifacts
  • Nominated reviewers with field expertise

Design principles

A benchmark grounded in research practice.

Open-ended science

A task’s outcomes cannot be adequately evaluated with a single scalar score. The scientific question leaves room for alternative approaches, discoveries, and expert interpretation.

Novel, frontier-relevant questions

Scientists propose problems with significance for their field. Review considers whether the direction is scientifically meaningful and has the potential for publication in a Nature-family journal.

Evaluation that learns from experiments

Rubrics begin as scientists’ best current criteria. Research outcomes reveal weaknesses that those criteria missed, prompting additions, revisions, and more informative experiments.

One interface for collaboration

OpenReview connects scientists and AI researchers. Scientists focus on questions and judgments; task managers and the harness handle execution and publish outcomes for review.

Have a scientific question to contribute?

Help define what good science means in your field.

Contribute a task

Looking for the supporting resources? Browse the research task catalog, catalog statistics, or tools and simulation library.