Skip to content

The science behind SciVeri-Bench

Can AI agents verify if AI agents are doing good science?

Our motivation, benchmark design, three scientific verification tasks, and scientist-in-the-loop collection process.

10 min readPublic documentationUpdated
On this page

Young-Jun Lee, Seungone Kim, Jinheon Baek, Soojung Yang, Yoonho Lee, Dongyeop Kang

Through this work, we aim to answer the question:

Can AI agents evaluate what constitutes good science as well as scientists can?

Current agent benchmarks rely on static evaluation#

Until last year, benchmarks based on competition problems, such as IMO, focused on how well reasoning models could solve given problems through deep reasoning, and performance on many of these benchmarks had largely saturated. As the agentic AI market has grown, a large number of agent benchmarks have emerged, including long-horizon agent benchmarks such as GAIA and tau3-bench; agentic coding benchmarks such as SWE-Bench Verified, TerminalBench 2.1, DeepSWE, and FrontierCode; and agentic computer-use benchmarks such as OSWorld 2.0. These recent benchmarks all evaluate agent capabilities in a static way. Some use deterministically verifiable scores as their headline metrics, while others use fixed rubrics created by domain experts and score outcomes through an LLM-as-a-Judge approach.

Science is alive, yet its evaluation remains static.#

More recently, the rise of AI for Science has led to many proposed science benchmarks, but these also continue to adopt static evaluation methods. Examples include PaperBench and NatureBench for paper reproduction; MLE-Bench for AI/ML engineering; FrontierResearch for proposing follow-up research directions; and LifeSciBench, TerminalBench-Science, and FrontierPhysics for expert-curated scientific problems. Nevertheless, we believe that conventional static, fixed-target evaluation has limitations when assessing whether AI agents are doing “good science” in the way domain scientists would judge it.

  1. Rubrics are not fully articulated until scientific experiments are conducted. Scientists do not necessarily create effective rubrics from the outset. Before running an experiment, it is difficult to judge which findings are truly valuable and which directions are unpromising. It is therefore inappropriate to keep using a fixed rubric established when the task was proposed—that is, before the experiment was run.
  2. AI agents do not know where the results fall short. Because AI agents do not know which aspects of experimental results are lacking, it is difficult for them to determine how to improve the experiments or generate rubrics. This is especially challenging for open-ended tasks, where the correct answer is unknown.
  3. AI agents should be able to judge which experimental results are good: Scientists draw on their experience and intuition to make these judgments, but AI agents find them difficult. Such judgment is necessary for refining the design of subsequent experiments and pursuing novel discoveries.

Therefore, as experiments progress, rubrics should evolve to reflect newly identified weaknesses, and AI agents should possess verification capabilities comparable to those of domain scientists. We believe this could be a starting point for the next stage of progress.

Contributions#

  • We propose the Scientist-Verified Benchmark (SciVeri-Bench), a benchmark for evaluating AI agents’ ability to perform scientific verification as human domain scientists do.
  • Our benchmark evaluates AI agents’ capabilities on three tasks: scientist rubric generation, scientist critique generation, and scientist taste evaluation.

🌟🌟🌟 All intellectual property rights to the outcomes generated by the agent for a proposed task will belong to the task proposer and reviewers. We will not claim ownership of these outcomes. You may use them to submit a paper to a Nature-family journal, and you are not required to include members of the co-lead team as co-authors.

SciVeri-Bench: Benchmarking AI Agents for Scientific Verification#

Guidelines for Task Collection: [Guideline] Track 4: SciVeri-Bench Task Proposal

Benchmark Design Principle#

One Interface for Collaboration. To create open-ended and novel scientific tasks, we consider the following points:

  1. Without help from domain scientists, it is hard to create open-ended and novel tasks. In addition, without multiple rounds of proposing and reviewing tasks, it will be difficult to come up with good open-ended and novel tasks.
  2. Even if domain scientists initially propose a task, it is very challenging for them to develop a good rubric for evaluating what constitutes “good” science without running experiments. This is because an understanding of what constitutes “good” science is learned through the experience of running scientific experiments.
  3. Domain scientists are not familiar with using AI agents or implementing code. Therefore, if we place too much responsibility for running agents on scientists, it may become a burden and make it difficult to obtain high-quality tasks. We need to enable domain scientists to do their best and concentrate on creating good open-ended and novel tasks that pose important scientific questions.

Therefore, we believe that when collecting scientific tasks from domain scientists, we need to provide an environment where AI researchers and domain scientists can interact while each group carries out its role as effectively as possible in its own setting. For example, AI researchers can run AI agents on proposed tasks in a computing environment, while domain scientists can think about and propose good scientific tasks wherever they feel comfortable, such as their lab or home. The two groups communicate through a single interface, allowing each to focus on its respective role. For the interface, we select OpenReview.

Scientists familiar with AI agents might run agents in their own environments and focus only on obtaining low scores (= challenging tasks), potentially at the expense of proposing truly good open-ended tasks. To prevent this, we separate the proposer and verifier, similar to the sandbox environment in Harbor. Scientists serve in two roles: proposers and verifiers. Their involvement is limited to proposing tasks, verifying the AI agent’s outputs, and refining their tasks and evaluation rubrics based on feedback from the executor (e.g., AI-generated outputs). These interactions take place through OpenReview, as illustrated in the figure below.

Full-size figure
Scientists and AI researchers collaborate through OpenReview.
Scientists and AI researchers collaborate through OpenReview.

Open-Ended and Novel Tasks. When collecting tasks for our benchmark, we use the following minimum criteria to define “open-ended” and “novel.” An open-ended task is one whose outcomes cannot be adequately evaluated using a single scalar score. For novelty and frontier relevance, we consider whether the task itself is significant enough to merit publication in a Nature-family journal.

Verification Tasks in SciVeri-Bench#

Full-size figure
The three scientific verification tasks in SciVeri-Bench.
The three scientific verification tasks in SciVeri-Bench.
  • Scientist Critique Generation: Given an AI outcome, can an evaluator identify its weaknesses?
  • Scientist Rubric Generation: Given an AI outcome and an existing rubric, can an evaluator identify problems in the outcome that the rubric does not yet cover? If a new problem appears in the AI outcome, can the evaluator turn it into a rubric criterion?
  • Scientist Taste Evaluation: Given 2~3 AI outcomes, can an evaluator judge which one represents better science than the others?

Task Collection Procedure#

Full-size figure
Task proposal, proposer-nominated review, and independent external review.
Task proposal, proposer-nominated review, and independent external review.

As shown in the figure above, our benchmark consists of two phases: (1) Phase 1: Task Proposal and Proposer-Nominated Review and (2) Phase 2: Independent External Review.

  • Phase 1: The task proposer nominates a reviewer from a similar scientific research field. This reviewer may be the proposer’s supervisor or collaborator.
  • Phase 2: An external reviewer who is unfamiliar with the task proposer and the Phase 1 reviewer but works in a similar field evaluates whether the proposed task, rubric, and resulting scientific findings are interesting, novel, and valuable enough for inclusion in SciVeri-Bench.

Note that Phase 2 will begin after mid-October.

Scientist-in-the-Loop Harness#

We collect open-ended and novel tasks from domain scientists through a “scientist-in-the-loop” harness. Our scientist-in-the-loop harness is fully automated for AI researchers and provides an environment where domain scientists can focus entirely on creating good tasks. Moreover, all communication between domain scientists and AI researchers can take place on OpenReview. In other words, scientists only need to check SciVeri-Bench 2026 Internal Review.

When collecting tasks, it is essential to solicit open-ended tasks whose outcomes cannot be adequately evaluated using a single scalar score. The problems addressed by these tasks must also be novel and at the research frontier.

What qualifies as “Novel & Frontier”? Whether the task itself is significant enough to merit publication in a Nature-family journal.

We intend to give contributors flexibility when proposing tasks. Examples include a problem they would like an AI agent to solve, a problem they consider difficult or novel in their field, follow-up research on a paper published in a Nature-family journal or a top conference within the past two years, or ongoing research in their lab. However, every task must be both open-ended and novel.

The overview of our “scientist-in-the-loop harness” is provided below.

Full-size figure
The scientist-in-the-loop harness and its iterative research process.
The scientist-in-the-loop harness and its iterative research process.

Step 1: Collecting Open-Ended Science Tasks through Multi-Round Peer Review#

The task proposer submits a task with the following information:

  • Task name
  • Task objective and description
  • Why the task is important
  • Data, software, and environment required to execute the task
    • If the task requires custom simulators or software developed in the proposer’s lab, the proposer must agree to make them publicly available.
    • If the task requires data that are not publicly available, the proposer must agree to make those data publicly available.
    • Agentic skills must not be included in the task submission.
  • An evaluation rubric defining what constitutes good science for the proposed task
  • Do you think there exists an ideal metric (e.g., a scalar score) that can well represent the quality of AI outcomes on your proposed task? If not, why?

The task proposer then nominates reviewers to evaluate the proposed task. The proposer may nominate more than two reviewers. Nominated reviewers should:

  • Have research experience in the scientific field relevant to the proposed task
  • Have one or two publications in Nature-family journals
  • Have prior experience conducting peer reviews for other journals or conferences

The nominated reviewers assess whether the proposed task is open-ended, novel, scientifically meaningful, and potentially suitable for publication in a Nature-family journal.

If the reviewers accept the task after multiple rounds of review, the task proposer implements the code for the proposed task and submits a ZIP file through OpenReview. If the proposer is unfamiliar with code implementation, the task manager helps implement the code. The task structure follows the Harbor format.

When implementing the task, specify the expected AI outputs in instruction.md. These may vary by task and could include task-specific artifacts and a paper.pdf.

Step 2: Executing AI Agents on the Proposed Task#

In this step, we run AI agents on the proposed task over multiple rounds to identify weaknesses in their outcomes, refine the evaluation rubric accordingly, and have the task proposer select their preferred AI outcome. This process forms another inner loop.

We use three frontier agents: Claude-Code Fable 5.1, Claude-Code Opus 5, and Codex GPT-6-Astra. In the first round, we run these agents on the proposed task to obtain AI outcomes, including task-specific artifacts and a paper. We then evaluate these outcomes using the initial rubric. If performance exceeds a certain threshold (to be determined empirically), we exit this inner loop and return to Inner Loop 1, where the task proposer and task reviewer modify the proposed task to make it more challenging based on the feedback.

Otherwise, we continue by showing the task proposer the outcomes from the three agents and conducting a survey with the following questions:

  • Are you satisfied with the current AI outcomes? If not, what do you consider their weaknesses? → For Scientist Critique Generation
  • What problems do you see, and how should they be reflected in the rubric? Should existing criteria be broken down, edited, or supplemented with new criteria? → For Scientist Rubric Generation
  • Among the two or three AI outcomes, which appears to have produced better findings or conducted more scientifically meaningful experiments? → For Scientist Taste Evaluation

When running the agents in the next round, we include the weaknesses identified in the previous round’s outcomes in instruction.md.

Once the scientist determines that no further work is needed, we send the final AI outcomes and rubric to the task reviewer, who assesses the quality of both.

Note that even if the iterative process saturates quickly, this is still acceptable, as our benchmark aims to understand how verification itself can be verified.

Independent External Review (TBD)#