Benchmarks
SaaS-Bench v1.1
🏆 Pine Computer / PCR · CS 78.3% · Computer-Use Agents
SaaS-Bench / v1.1
Can computer-use agents finish real SaaS workflows?
Results across 106 tasks in 23 deployable SaaS applications. The v1.1 table compares a shared, open-source browser harness with each model's native agent runtime.
78.3%Top CS · Pine Computer / PCR
31.1%Top RS · Opus 5 / Claude Code
106Tasks · 6 domains
v1.1 Results
CS = Checkpoint Score, measuring partial progress across verification checkpoints. RS = Resolved Score, the percentage of tasks completed end to end. The same model can have different results under different harnesses, so each model–harness pair is shown separately.
| Model | Harness | Step / task | Cost / task | Tool Call / task | CS (%) | RS (%) |
|---|---|---|---|---|---|---|
| Generic Harness | ||||||
| GPT-5.6 sol | Browser-Use (Open-Source) | 142.7 | ~$14.9 | 257.7 | 69.8 | 17.9 |
| Opus 5 | Browser-Use (Open-Source) | 162.1 | ~$20.0 | 239.4 | 64.7 | 21.7 |
| Qwen 3.8 max | Browser-Use (Open-Source) | 71.5 | ~$2.5 | 131.3 | 31.1 | 8.5 |
| Kimi K3 | Browser-Use (Open-Source) | 121.9 | ~$6.8 | 237.8 | 56.2 | 17.0 |
| Native Harness | ||||||
| GPT-5.6 sol | Codex | 250.7 | ~$20.5 | 250.7 | 71.1 | 29.2 |
| Opus 5 | Claude Code | 270.5 | ~$26.5 | 270.5 | 74.3 | 31.1 |
| Qwen 3.8 max | Qwen Code | 268.9 | ~$5.1 | 268.9 | 65.9 | 23.6 |
| Kimi K3 | Kimi-CLI | 284.4 | ~$9.4 | 284.4 | 56.4 | 24.5 |
| Pine Computer | Pine Computer Runtime (PCR) | 335.5 | ~$3.6 | 306.2 | 78.3 | 27.4 |
Values are from the v1.1 results table. Cost per task is approximate, in USD. CS and RS are percentages. On narrow screens, scroll the table horizontally to view every column.
What the results show. Under the generic harness, GPT-5.6 sol has the highest CS (69.8%), while Opus 5 has the highest RS (21.7%). Under native harnesses, Pine Computer / PCR has the highest CS (78.3%), while Opus 5 / Claude Code has the highest RS (31.1%).