PaperBenchX
93 paper-reproduction tasks across 12 research areas, scored through clean replay and evidence-backed scientific rubrics.
93 paper-reproduction tasks across 12 research areas, scored through clean replay and evidence-backed scientific rubrics.
v1.1 results across 106 tasks and 23 SaaS systems, comparing generic and native harnesses using Checkpoint Score and Resolved Score.
A continuously updated benchmark evaluating AI coding agents on real-world software engineering tasks from GitHub issues.
115 real-world version-upgrade tasks from 17 repositories across 5 programming languages. Evaluates whether coding agents can implement substantial new functionality guided only by a multi-target specification.
26 multi-turn coding tasks with 227 evaluated rounds. Each task keeps the same workspace and agent session while requirements change, accumulate, and sometimes conflict.
Can AI browser agents complete everyday online tasks on live websites? ClawBench evaluates agents across V1/V2 tasks with isolated Chrome runs, request interception, five-layer traces, and agentic judging.
A dynamic evaluation engine for AI prediction systems, featuring multi-point aligned Elo ranking, three-track data collection, and adaptive scheduling across diverse domains including finance, politics, crypto, sports, and esports.
A benchmark for visual reasoning that challanges frontier MLLMs yet 3-year-olds can solve.