Benchmarks

PaperBenchX

GPT-6-Astra / Full 13.98% · Scientific Reproduction

PaperBenchX / Scientific reproduction

Can AI agents reproduce published scientific results?

PaperBenchX evaluates whether research agents can reconstruct experiments from papers, operate domain-native scientific software, and regenerate evidence that supports the original claims. The benchmark contains 93 tasks across 12 research areas and six scientific domains.

13.98%Top Full rate / 13 of 93 tasks
59.08%Top Partial score / GPT-6-Astra
3,168Scorable criteria / 12 research areas

Overall results

Full is the percentage of tasks scoring strictly above 90%. Partial averages rubric credit over the fixed 93-task set. Stage scores show where credit is earned across Modeling, Execution, and Validation; average turns describe interaction effort. Each rubric item is evaluated by either a deterministic check or an LLM judge. Items requiring scientific judgment use GPT-5.6-Sol as the shared judge across all agent configurations.

AgentFull (%)Partial (%)ModelingExecutionValidationAvg turns
GPT-6-Astra (Codex)13.9859.0872.7269.3552.76348.7
Fable 5.112.9058.1771.0069.9251.06113.0
Fable 511.8357.9672.1368.9950.39112.0
GPT-5.6-Sol (Codex)9.6855.0769.8571.2345.29237.8
Opus 57.5352.5266.9967.1647.76259.2
Kimi-K37.5349.2462.4060.7042.32185.9
Qwen3.8-Max5.3846.7664.4161.0241.65234.4
DeepSeek-V4.1-Flash7.5345.4759.8757.5141.13372.6
GLM-5.37.5340.7851.1151.8537.06290.8
HY32.1527.5336.8340.5919.67197.4
What the results show. The leading system earns substantial partial credit but fully reproduces only 13 of 93 tasks. Modeling and execution are stronger than validation across the cohort, showing that running scientific software is not enough: agents must also produce replayable evidence that supports the intended scientific claim.