For task contributors
Track 4: SciVeri-Bench Task Proposal
The detailed guide for scientists proposing a research task: what to prepare, how to submit, and how scientific review works.
On this page
Project Lead: Young-Jun Lee (Contact: [email protected])
If you did not select this track, you do not need to complete it. However, if you would like to participate even though you did not indicate your intention to do so on the Google Form, please let us know via the Slack channel.
👋 Welcome SciVeri-Bench Task Proposal Track!#
This page provides step-by-step guidance on how to submit your proposed task and interact with the reviewer and manager. It also explains the overall workflow of our scientist-in-the-loop harness. If you are unfamiliar with the “scientist-in-the-loop” harness, please read the overview here before participating in this track.
Prerequisite#
- Task proposers and reviewers should have an OpenReview profile. We will invite you as Task Proposer or Task Reviewer in our venue through your institution email. If you did not receive this email, please let us know in Slack.
- Task proposer should nominate a reviewer on their proposed task though Nominate a Task Reviewer for SciVeri-Bench
Eligibility for Task Proposers and Task Reviewers#
- Task proposers should have published 1–2 papers in Nature-family journals or at top-tier conferences within the past two years.
- Nominated task reviewers should have published 1–2 papers in Nature-family journals or at top-tier conferences. AI conferences such as NeurIPS, ICLR, ICML, and CVPR are acceptable. However, please note that in Phase 2, tasks will be evaluated by external reviewers who have published in Nature-family journals.
- Nominated task reviewers should have completed 1–2 or more peer reviews for Nature-family journals or top-tier conferences.
Task Submission Deadline#
- There is no fixed submission deadline. The goal is to develop high-quality, open-ended, and novel tasks, along with strong verification materials, through active and ongoing interaction between task proposers and task reviewers.
🚨 Exception: Contributors to SciVeri-Bench v0.1 must complete their tasks within 2–3 weeks. I will notify the relevant contributors on Slack.
🍀 How to Contribute Your Task#
Our task collection procedure is based on our proposed “scientist-in-the-loop” harness. Please follow the step-by-step instructions below.
Step 0: Propose an Open-Ended and Novel Task#
When proposing a scientific task, ensure that it is open-ended and novel by considering the following minimum requirements:
- The task must be open-ended, with outcomes that cannot be adequately evaluated using a single scalar score
- There must be no reference solution available before the experiment is conducted (i.e., because the task is open-ended, you do not know in advance what will constitute a good or poor discovery)
- Even if a task is open-ended, it must not be a reproduction of a previously published paper
- The problem addressed by the task must be novel and at the research frontier
- The task itself must be significant enough to merit publication in a Nature-family journal
🌟 All intellectual property rights to the outcomes generated by the agent for a proposed task will belong to the task proposer and reviewers. We will not claim ownership of these outcomes. You may use them to submit a paper to a Nature-family journal, and you are not required to include members of the co-lead team as co-authors.
These are only the minimum requirements for an open-ended and novel task. We aim to give contributors flexibility in proposing tasks. Examples include a problem they would like an AI agent to solve, a problem they consider difficult or novel in their field, follow-up research on a paper published in a Nature-family journal or at a top conference within the past two years, or ongoing research in their lab. However, every task must be both open-ended and novel.
The requirement that tasks cannot be evaluated solely using a single scalar score does not exclude score-based optimization tasks, such as improving a black-box model’s score beyond that of existing models. Such tasks are welcome if their evaluation also involves understanding and analyzing novel discoveries or elucidating mechanisms revealed through the optimization process, which are typically required to use rubric-based metrics (not only a single scalar score).
How to propose your task? Using LaTeX file!#
When proposing a task, use the provided LaTeX file (Overleaf, Online LaTeX Editor) to describe your submission and include the following information:
- Task Name
- Task Objective
- Why it matters
- Data, software, and environment required to execute the task
- If the task requires custom simulators or software developed in the proposer’s lab, the proposer must agree to make them publicly available.
- If the task requires data that are not publicly available, the proposer must agree to make those data publicly available.
- Evaluation
- An evaluation rubric defining what constitutes good science for the proposed task (required)
- Not all rubric criteria award positive scores; some assign negative scores to penalize actions the AI agent should not attempt or discoveries it should not make
- score-based evaluation metric
- An evaluation rubric defining what constitutes good science for the proposed task (required)
- Deliverables (What to hand in)
- A paper in PDF format (required)
- Task-specific artifacts (optional)
The provided LaTeX file includes a template for the information above. You are welcome to add any other information you consider necessary or helpful.
- You are free to organize your submission in any format. However, the five main sections—Task Objective, Why it matters, Data, software, and environment required to execute the task, Evaluation, and Deliverables—must be included with these exact titles. You may add other sections anywhere in the document.
- Within these five main sections, you may freely organize subsections and arrange tables and figures.
Step 1: Access into SciVeri-Bench Venue in OpenReview#
- Log into OpenReview
- Access into our venue OpenReview: SciVeri-Bench 2026 Internal Review
Step 2: Submit Your Task#
- Click “SciVeri-Bench 2026 Internal Review Submission”
-
Fill in the following information:
Information
- Task Name: The full, descriptive title of the proposed task
- Task Identifier: A short identifier for the task in snake_case (e.g., materials_discovery)
- Authors: Your name
- Domain: Select the primary scientific domain of the proposed task
- Domain Other: If you selected Other as the primary domain, please specify the scientific domain here
- Sub Domains: Comma separated list of science sub-domains of the proposed task
- Keywords: Comma separated list of keywords of your proposed task
- Task Description: Brief description of task. (since a detailed task description is provided in the LaTeX PDF)
- Evaluation Metric Adequacy: Do you think there is an ideal metric (e.g., a scalar score) that can adequately represent the quality of AI-generated outcomes on your proposed task? If not, why?
- Software Release Agreement: If the task requires custom simulators or software developed in your lab, you must agree to make them publicly available
- Data Release Agreement: If the task requires data that are not publicly available, you must agree to make those data publicly available
- PDF: Upload your PDF file where it contains what your proposed task is.
- Supplementary Material: Upload any supplementary materials as a ZIP archive
- Supplementary Material Link: Provide a link to supplementary materials hosted on Google Drive or another file-sharing service, especially if they exceed the upload limit. Ensure reviewers and task managers can access the files or folder without requesting additional permission
- Data Release: If your task submission is accepted following internal review, the task submission, author names, and the review and discussion history in the submission thread, including exchanges with reviewers, will be made publicly available. Please confirm that you agree to this release
- License: Select CC-By-4.0
🚨 The task proposer must not discuss the task with the nominated reviewer before the multi-round peer review begins.
Step 3: Start Multi-round Peer Review#
Task Reviewer
- First, read the PDF file submitted by the task proposer. Then, assess whether the proposed task is novel, scientifically meaningful, and potentially suitable for publication in a journal in the Nature family.
- The reviewer should provide a score for the proposed task at every turn (i.e., each time they leave a comment).
- After multiple rounds of discussion with the task proposer, if the reviewer considers the proposed task ready to proceed, they should provide a final score and leave a comment. This comment must include the sentence “It is acceptable.”
Task Proposer
- Based on the task reviewer’s comments, please refine your proposed task to make it more novel and scientifically meaningful.
End Condition
- When the task reviewer says “It is acceptable”, the task proposer moves on to the next step.
Step 4: Implement Your Proposed Task#
Task proposers should implement their code in Harbor format. At this stage, the Task Manager will help ensure that the code follows the Harbor format and works correctly. If task proposers are unfamiliar with coding, the Task Manager should provide the support they need to successfully implement their code.
After completing the code, the task proposer should click the “Task Upload” button and upload a ZIP file containing the code to the “Task Zip” field. They must then enter /submit_task in the “Comment” field, as shown in the figure below.
After submission, the request waits for approval. The assigned Task Manager must reply directly to the original /submit_task request using Official Comment, calling only /validate_task in the Comment field.
After approval, the scientist-in-the-loop harness automatically performs two checks:
- Task Package Validation: Checks the ZIP structure, required task files, and Harbor configuration.
- Environment Validation: Builds the environment from the submitted Dockerfile, or loads a configured prebuilt image, and verifies container startup, configured health checks, and basic command execution.
If either check fails, the harness automatically posts an alert on the original request in OpenReview, identifying the failed stage and the error details. Temporary infrastructure errors may be retried before an alert is posted.
If all checks pass, the harness automatically posts a “Task preflight validation passed” message in the same OpenReview thread, as shown in the image below.
Step 5: Execute Agents#
After all checks pass, then Task Manager should input /run_agent command in Comment to execute three frontier agents, Codex (GPT-6-Astra), Claude-Code (Opus 5.5), and OpenScience (GPT-6-Astra). Task Manager can customize the command such as
/run_agent --agent codex --model gpt-6-astra --effort max/run_agent codex






