AI research agent working across scientific computing environments

PaperBenchX: Can AI Reproduce Science Across Disciplines?

GitHub

01 / Overview

Can AI reproduce published scientific results?

Reproducibility is a defining feature of science. Before an AI system can be trusted to discover something new, it should be able to reconstruct and verify what science already knows. PaperBenchX tests this demanding middle ground: the target claim is known, but the agent must turn an incomplete scientific description into an executable experiment and produce evidence that another researcher can independently regenerate.

PaperBenchX contains 93 tasks across twelve research areas in six scientific domains
93 paper-reproduction tasks across 12 research areas and six scientific domains.
93Paper-reproduction
tasks
12Research
areas
10Domain-native
environments
3,168Scorable rubric
criteria

The tasks draw on 93 papers from 44 publication venues. Agent time limits range from 4 to 24 hours, with a median of 7 hours.

02 / Why scientific reproduction?

A common interface does not make the science uniform.

PaperBenchX gives every task the same outer contract: a paper, public instructions, a configured environment, and a runnable submission entry point. That standardizes how an agent starts and how its work is delivered. It does not supply the missing scientific decisions needed to reproduce the result.

Inside that shared interface, the work changes substantially across HFSS, Meep, electronic-structure codes, data-analysis pipelines, and embodied simulators. Geometry, boundary conditions, physical models, numerical tolerances, convergence behavior, and result extraction determine what experiment was actually performed. A command can finish successfully and still model the wrong system or support the wrong claim.

01 / ModelingWhat experiment did the paper perform?Translate prose, equations, and figures into the correct model, assumptions, parameters, and target observable.
02 / ExecutionDid the scientific computation work?Operate specialized software, interpret diagnostics, and determine whether a run is complete, stable, and converged.
03 / ValidationDoes the evidence support the claim?Regenerate the required artifacts and verify that they arise from the intended workflow rather than a plausible-looking shortcut.

The benchmark evaluates three connected stages: Modeling recovers the intended experiment, Execution runs it in its native environment, and Validation tests the regenerated evidence. Passing only one stage is not a successful reproduction.

03 / Inside a task

What does a PaperBenchX task look like?

Scientific correctness cannot be established from an agent's final report alone. PaperBenchX therefore combines clean replay, evidence-based grading, and hierarchical scoring to test whether a submitted procedure regenerates evidence that supports the target claim—not merely a plausible-looking number.

PaperBenchX task inputs, agent submission, clean replay, and private evidence-based evaluation
Task inputs → agent submission → clean replay → rubric evaluation. The Brewster criteria, pass/fail marks, and 24.5% outcome illustrate grading, not a measured result or a universal pass rule. Platform logos are examples, not a complete inventory. Click to enlarge.
  1. Task inputs — define what must be reproduced

    The paper, instructions, and addendum specify the scientific target, required methods, permitted alternatives, and evidence to deliver. Private grading must assess stated obligations, not introduce hidden requirements.

  2. Agent submission — hand over the experiment

    The agent submits source code, configuration, permitted inputs, and a runnable reproduce.sh. The package must regenerate the required outputs; a report or a folder of saved figures cannot replace the workflow that produced them.

  3. Clean replay — establish where the evidence came from

    Remove generated outputs and stale state, then run the submitted workflow again. The protocol requires isolation, blocked network access, and task-specific limits. Admissible outputs are frozen for grading, so surviving files from the agent's earlier session do not count as regenerated evidence. Replay regenerates the evidence; rubric evaluation tests whether it supports the scientific claim.

  4. Rubric evaluation — test each scientific claim

    Each rubric item is evaluated using one of two methods: deterministic checks for directly verifiable numerical or structural properties, or LLM-based evaluation for items requiring scientific interpretation and judgment. GPT-5.6-Sol serves as the shared LLM judge across all agent configurations.

    Rubrics and scoring

    Illustrative rubric tree: evidence-backed criterion scores between 0 and 1 combine through requirement groups into task score R = Σ wᵢsᵢ. Effective weights reflect scientific importance and sum to 1; branches and weights vary by task.
    Illustrative tree · task-specific weights
    Full rate R > 0.90
    Fraction of the 93 tasks strictly above the threshold.A score of exactly 0.90 does not qualify.
    Partial score
    Mean task score across the same fixed set of 93 tasks.Missing outcomes count as zero.
04 / Task construction

From papers to reproduction tasks

Turning a paper into a benchmark task requires more than extracting a target number. Each task begins with a verifiable scientific claim and passes through two stages. In Build, a domain expert selects the claim, fixes the scope and conditions, provisions the native environment, writes the public task contract, and completes a private reference workflow. In Validate and Review, the expert defines evidence-backed rubric leaves, reruns the reference workflow from a clean environment, calibrates the scoring on valid and invalid submissions, and sends the task through independent review.

The PaperBenchX two-stage task construction pipeline, from selecting a published claim to clean reference replay, calibration and independent expert review
Build, then validate and review. Tasks that fail because of scope, environment, or reference execution return to the build stage; scoring problems return to rubric construction. Two independent domain experts review the scientific scope, fairness, reproducibility, evidence requirements, and scoring behavior before a task is frozen.

This process produced 93 tasks with 3,168 scorable criteria, a median of 38 per task. The environment is part of the scientific task itself: model representations, numerical conventions, convergence evidence, and execution procedures differ across software.

05 / Leaderboard

How do current agents compare?

Across ten frontier agent configurations, the strongest system fully reproduces only 13 of 93 tasks under the strict 90% threshold, a 13.98% Full rate. Its Partial score reaches 59.08%, but the large gap between partial and full completion shows that substantial progress rarely closes the full loop from paper understanding to reproducible scientific evidence.

13.98%Best Full rate
13 of 93 tasks
59.08%Best Partial score
GPT-6-Astra
71Tasks never fully
reproduced

Overall leaderboard

Every configuration is evaluated on the same fixed set of 93 tasks.

10 agents · 93 tasks
Full Scorex.x%Partial Score
Performance (%)

Hover, focus, or select a model to compare its Full and Partial scores.

Bars show Full scores (>90%); boxes show Partial scores.
Loading the leaderboard…
How to read these rankings

Full is the fraction of tasks scoring strictly above 90%; a score of exactly 90% does not qualify. Partial averages rubric scores over the fixed 93-task set and captures supported progress on incomplete reproductions. Missing outcomes count as zero in both views.

Every configuration is compared on the same task denominator. Ties use unrounded scores, and all submissions are evaluated through clean replay with the same scientific rubric evaluator and deterministic checks.

Research-area difficulty

Hardest to easiest for the evaluated ten-model cohort

12 research areas

Select a research area to inspect its cross-model mean Partial score.

Difficulty = 100 − cross-model mean Partial score. This is a comparative result for the evaluated cohort, not an intrinsic property of the papers.

06 / Analysis

What separates partial progress from reliable reproduction?

The overall score shows who performs better; the three views below explain where the difference comes from. Read them together, then hover or focus on a mark for its exact value.

Three questions behind the score

Stage performance, interaction efficiency, and workflow allocation

Stage performance, interaction length, and workflow allocation.
01

The main drop appears at scientific validation.

Modeling and Execution reach median scores of 73.4% and 73.5%, while Validation falls to 40.0%. Even among 304 outcomes with at least 80% in both upstream stages, only 53 cross the full-reproduction threshold. Agents often build and run much of an experiment without closing the evidentiary loop.

02

More interaction does not reliably produce a better result.

Fable 5.1 and Fable 5 reach about 58% Partial score in roughly 113 turns. GPT-6-Astra reaches 59.08% in 348.7 turns, while DeepSeek-V4.1-Flash uses 372.6 turns for 45.47%. The weakly negative correlation (r = −0.14) suggests that long trajectories often reflect recovery attempts or continued work on a mis-specified experiment.

03

Most time is spent executing, but success is decided upstream.

Native software consumes 51.8–63.3% of tool-mediated time. On the F-stub antenna task, however, Fable 5 recovers more of the required scientific evidence in 159 turns than GPT-6-Astra or HY3 do in 464 and 390 turns. Correct modeling, parameter choices, and intermediate checks determine whether expensive computation becomes useful evidence.

04

Stage-aware scoring distinguishes different kinds of failure.

A run can complete yet support the wrong conclusion because its modes, ports, parameters, mechanism, or result extraction are incorrect. The LNOI edge-coupler result illustrates this separation: GPT-6-Astra receives full Execution credit but only 41.9% in Validation. Clean replay and stage scores distinguish an execution failure from an unsupported scientific claim.

07 / Data release

12 open tasks, 81 held-out tasks

PaperBenchX releases one representative task from each research area: 12 open tasks that the community can run, inspect, and use to study the strengths and limitations of research agents. The remaining 81 tasks are maintained as a held-out evaluation set with controlled access, limiting benchmark contamination while preserving long-term measurement value.

12

Open tasks

Run the tasks, inspect the evaluation method, debug reproduction workflows, and examine how individual scientific requirements become scores.

81

Held-out tasks

Retained for continuing evaluation of future agents, helping PaperBenchX preserve discrimination as systems improve.

What the open release provides

The release includes task instructions, scoring details, and the resources required for evaluation. Scientific software licensing determines how each environment is distributed:

9 tasks

Redistributable environments

Complete open-source scientific computing environments are provided and can be run directly.

3 tasks

Licensed environments

For tasks that depend on commercial software, we release the task instructions, scoring details, environment setup guidance, and unified evaluation framework. Researchers can configure and run the environment after obtaining a valid license.

08 / Toward a general AI scientist

Reliable reproduction is a foundation for new science.

PaperBenchX makes that foundation measurable. Agents already make partial progress in Modeling and Execution, yet the strongest configuration fully reproduces only 13.98% of tasks. A substantial gap remains between reading a paper, operating scientific software, and regenerating evidence that supports its claims.

Measured by PaperBenchX01

Modeling

Recover the intended model, assumptions, and experiment.

Measured by PaperBenchX02

Execution

Operate domain-native tools and diagnose the computation.

Measured by PaperBenchX03

Validation

Regenerate evidence and test whether it supports the claim.

Beyond reproduction04

Form hypotheses

Identify open questions and prioritize ideas worth testing.

Beyond reproduction05

Discover

Design informative tests and revise beliefs with new evidence.

What PaperBenchX contributes

A measurable first step

The benchmark turns gaps in cross-disciplinary reproduction into evidence that can guide improvement. Reproduction is not discovery, but it provides a verifiable checkpoint on the path toward it.

What lies beyond

From computation to physical feedback

PaperBenchX currently focuses on scientific tasks that can be completed on a computer. Extending this direction to wet-lab research would require agents to design experiments, act safely in the physical world, learn from observations, and revise hypotheses when evidence disagrees.

PaperBenchX makes reliable scientific reproduction a measurable step toward more autonomous scientific research.

09 / Contributors

Contributors

Core contributorsRuoyu Wu1, Zengji Tu1, Aolong Sun1, Tianyi Ma1, Jin Chen2, Shen Yan2, Liang Chen1, Kuan Li1

ContributorsJunren Li1, Li Fu1, Ziyan Zhang1, Pinrui Huang1, Xinbo Xu1, Xiaochen Wang1, Xu Li2, Zhongfei Hou2, Nan Shu2, Chengda Wu1, Yutao Zhou1, Chaoyi Huang2, Sihang Yuan2, Zihan Xu2

AdvisorsBaobao Chang3, Lin Chang3, Jianjun Wu3

1UniPat AI 2ByteDance Seed 3Peking University

For further collaboration on PaperBenchX, contact paperbenchx@unipat.ai.