PaperBenchX: Can AI Reproduce Science Across Disciplines?
PaperBenchX evaluates whether AI agents can reconstruct published scientific experiments, execute domain-native software, and regenerate verifiable evidence across disciplines.
PaperBenchX evaluates whether AI agents can reconstruct published scientific experiments, execute domain-native software, and regenerate verifiable evidence across disciplines.
AI agents can browse the web — but can they actually do your job? SaaS-Bench puts computer-use agents inside 23 real SaaS systems to find out, with 106 professional workflows spanning finance, healthcare, engineering, and more.
UniSwarm coordinates research-agent swarms through a unified interface and complements parallel thinking as a new axis of test-time scaling. We introduce UniSwarm-35B-A3B, trained on UniScientist data to act as both the main agent and a sub-agent, achieving competitive performance against frontier research agents.
Most existing agent benchmarks evaluate only the outcome of a single execution—whether the final result passes or fails, but this fundamentally differs from how people actually use coding agents. Vibe coding is iterative and human-in-the-loop: users inspect intermediate results, refine requirements, correct mistakes, and steer agents across multiple rounds. We introduce Vibe-Coding Arena, an evaluation system that captures this process by having humans build alongside agents and conduct fine-grained evaluations of their behavior throughout the interaction.
UniMath-35B-A3B is an open-source olympiad gold-medal mathematical model post-trained from Qwen3.6-35B-A3B. Fine-grained proof-evolution data activates its self-evolving reasoning capability, enabling human gold-medal-level performance on IMO 2025 and USAMO 2026.
ExpertEval is a large-scale, expert-annotated evaluation infrastructure spanning Medicine, Finance, and Law. We measure productive intelligence: the capacity to reason under genuine professional constraints where errors carry irreversible consequences.
Terminal-X evaluates coding agents in executable terminal tasks. It combines DeepTerminalBench for single-shot depth, EvoCode-Bench for multi-turn iteration, and RoadmapBench for version-upgrade work on real repositories.
We present Echo, a full-stack prediction intelligence system centred on EchoZ-1.0, the first large language model trained end-to-end under the Train-on-Future paradigm — spanning a dynamic evaluation engine, a post-training pipeline, and an AI-native prediction API.
While coding capabilities have surpassed human-level performance in many benchmarks, visual reasoning continues to lag behind. In this work, we introduce SWE-Vision, a minimal agentic workflow that leverages a simple coding environment to enhance visual understanding, also a more achievable test time scaling direction.
UniScientist is designed to advance universal scientific research intelligence through a unified paradigm. Leveraging an evolving polymathic synthesis, we generate research-grade data that enables structured, rubric-based supervision.
State-of-the-art MLLMs achieve PhD-level language reasoning but struggle with visual tasks that 3-year-olds solve effortlessly. We introduce BabyVision, a benchmark revealing the infancy of AI vision.