Benchmarks

SaaS-Bench v1.1

🏆 Pine Computer / PCR · CS 78.3% · Computer-Use Agents

SaaS-Bench / v1.1

Can computer-use agents finish real SaaS workflows?

Results across 106 tasks in 23 deployable SaaS applications. The v1.1 table compares a shared, open-source browser harness with each model's native agent runtime.

78.3%Top CS · Pine Computer / PCR
31.1%Top RS · Opus 5 / Claude Code
106Tasks · 6 domains

v1.1 Results

CS = Checkpoint Score, measuring partial progress across verification checkpoints. RS = Resolved Score, the percentage of tasks completed end to end. The same model can have different results under different harnesses, so each model–harness pair is shown separately.

ModelHarnessStep / taskCost / taskTool Call / taskCS (%)RS (%)
Generic Harness
GPT-5.6 solBrowser-Use (Open-Source)142.7~$14.9257.769.817.9
Opus 5Browser-Use (Open-Source)162.1~$20.0239.464.721.7
Qwen 3.8 maxBrowser-Use (Open-Source)71.5~$2.5131.331.18.5
Kimi K3Browser-Use (Open-Source)121.9~$6.8237.856.217.0
Native Harness
GPT-5.6 solCodex250.7~$20.5250.771.129.2
Opus 5Claude Code270.5~$26.5270.574.331.1
Qwen 3.8 maxQwen Code268.9~$5.1268.965.923.6
Kimi K3Kimi-CLI284.4~$9.4284.456.424.5
Pine ComputerPine Computer Runtime (PCR)335.5~$3.6306.278.327.4

Values are from the v1.1 results table. Cost per task is approximate, in USD. CS and RS are percentages. On narrow screens, scroll the table horizontally to view every column.

What the results show. Under the generic harness, GPT-5.6 sol has the highest CS (69.8%), while Opus 5 has the highest RS (21.7%). Under native harnesses, Pine Computer / PCR has the highest CS (78.3%), while Opus 5 / Claude Code has the highest RS (31.1%).