PaperBenchX
GPT-6-Astra / Full 13.98% · Scientific Reproduction
Can AI agents reproduce published scientific results?
PaperBenchX evaluates whether research agents can reconstruct experiments from papers, operate domain-native scientific software, and regenerate evidence that supports the original claims. The benchmark contains 93 tasks across 12 research areas and six scientific domains.
Overall results
Full is the percentage of tasks scoring strictly above 90%. Partial averages rubric credit over the fixed 93-task set. Stage scores show where credit is earned across Modeling, Execution, and Validation; average turns describe interaction effort. Each rubric item is evaluated by either a deterministic check or an LLM judge. Items requiring scientific judgment use GPT-5.6-Sol as the shared judge across all agent configurations.
| Agent | Full (%) | Partial (%) | Modeling | Execution | Validation | Avg turns |
|---|---|---|---|---|---|---|
| GPT-6-Astra (Codex) | 13.98 | 59.08 | 72.72 | 69.35 | 52.76 | 348.7 |
| Fable 5.1 | 12.90 | 58.17 | 71.00 | 69.92 | 51.06 | 113.0 |
| Fable 5 | 11.83 | 57.96 | 72.13 | 68.99 | 50.39 | 112.0 |
| GPT-5.6-Sol (Codex) | 9.68 | 55.07 | 69.85 | 71.23 | 45.29 | 237.8 |
| Opus 5 | 7.53 | 52.52 | 66.99 | 67.16 | 47.76 | 259.2 |
| Kimi-K3 | 7.53 | 49.24 | 62.40 | 60.70 | 42.32 | 185.9 |
| Qwen3.8-Max | 5.38 | 46.76 | 64.41 | 61.02 | 41.65 | 234.4 |
| DeepSeek-V4.1-Flash | 7.53 | 45.47 | 59.87 | 57.51 | 41.13 | 372.6 |
| GLM-5.3 | 7.53 | 40.78 | 51.11 | 51.85 | 37.06 | 290.8 |
| HY3 | 2.15 | 27.53 | 36.83 | 40.59 | 19.67 | 197.4 |