Leaderboard
Agent scope| System / Submission | Score | Organization | Reported | Source |
|---|---|---|---|---|
| Claude Fable 5.1 New Partial score (41.7% strict) on the 2026.08.08 release at 1080p, 500 steps, Opus 4.8 grader; the card states this run modified tasks and grading, so it is not directly comparable to official-settings rows. Self-reported. | 77.9% | Anthropic | Source | |
| Simular Sai New Partial score (28.25% binary) at $15.70 per task for Simular's neuro-symbolic agent; self-reported by Simular, a benchmark co-author, against its quoted Opus 5 and Sol comparators. | 73.0% | Simular AI | Source | |
| GPT-6 Astra New Partial score on the 2026.08.08 offline subset under a latency simulation at roughly 40 minutes per task; maximum at any effort. Self-reported at launch. | 72.6% | OpenAI | Source | |
| Claude Opus 5 New Partial score (five-run first-attempt average at 1080p, 500 steps, Opus 4.8 grader); OpenAI's launch table separately lists 70.2% for its author-reproduced run on the offline subset. Self-reported. | 70.6% | Anthropic | Source | |
| GPT-6.1 Sol New Partial score on the v2026.08.08 offline subset at maximum reasoning effort. Derived from the launch statement that Sol is 2.1 percentage points below Astra (72.6%); self-reported by OpenAI. Not directly comparable to combined online/offline scores. | 70.5% | OpenAI | Source | |
| Gemini 4 Argon New Partial score on the 2026.08.08 offline subset; best of three single-attempt runs at 1080p and 500 steps using the official evaluator, Gemini CUA harness, batched tools, and compaction. Self-reported by Google; not directly comparable to combined online/offline scores. | 69.2% | Source | ||
| Claude Opus 5 (Snorkel run) New Partial score (31.43% binary) in Snorkel's independent run at max effort with batched tools under the 500-step budget; Snorkel is a benchmark co-author. | 68.31% | Snorkel AI | Source | |
| Muse Spark 1.3 New Partial score (32.0% binary) at max reasoning effort on benchmark version 08.08 in Meta's GUI computer-control harness on a full Ubuntu desktop VM. Self-reported on the launch scorecard. | 66.9% | Meta | Source | |
| Claude Fable 5 New Partial score from the Opus 5 card's comparison table, which sources competitor scores from their release posts. Self-reported. | 66.1% | Anthropic | Source | |
| GPT-5.6 Sol (offline subset) New Partial score on the 2026.08.08 offline subset under a latency simulation at roughly 75 minutes per task; a separate launch-post figure from the Sol row below. Self-reported. | 65.7% | OpenAI | Source | |
| GPT-5.6 Sol (Snorkel run) New Partial score (27.34% binary) in Snorkel's independent run under the 500-step budget; Snorkel is a benchmark co-author. | 62.72% | Snorkel AI | Source | |
| GPT-5.6 Sol New OpenAI self-reported at the GPT-5.6 launch; single-agent Sol (ultra multi-agent not reported on OSWorld 2.0). Partial score; binary completion not published. | 62.6% | OpenAI | Source | |
| Claude Opus 4.8 (Opus 5 card table) New Partial score from the Opus 5 card's comparison table; the benchmark authors' official-settings run of the same model reaches 54.8% partial (20.6% binary). Self-reported. | 55.7% | Anthropic | Source | |
| Claude Opus 4.8 (batched tools) New Author-run on the official OSWorld 2.0 harness (108 long-horizon tasks); max thinking with batched tool calls, 500-step budget. 20.6% binary completion. | 54.8% | Anthropic | Source | |
| Qwen3.8-Flash-Next New Partial score (19.4% binary) under the card's partial/binary reporting; benchmark release version not stated. Self-reported on the model card. | 52.3% | Alibaba | Source | |
| GPT-5.5 (batched tools) New Author-run; xhigh reasoning with batched tool calls, 500-step budget (~$2,750 per run). 13.0% binary completion, flat across 150/300/500 steps. | 49.5% | OpenAI | Source | |
| Claude Opus 4.8 New Author-run; max thinking, standard tool calls, 500-step budget. 18.52% binary completion. Batched tool calls lift the same model to 54.8% partial (rank 2). | 49.33% | Anthropic | Source | |
| Claude Opus 4.7 New Author-run; max thinking, standard tool calls, 500-step budget (~$3,870 per run). 13.9% binary completion. | 49.1% | Anthropic | Source | |
| Claude Opus 4.7 (batched tools) New Author-run; max thinking with batched tool calls, 500-step budget. 18.2% binary completion. | 48.91% | Anthropic | Source | |
| Claude Sonnet 4.6 (max thinking) New Author-run; max thinking, standard tool calls, 500-step budget (~$2,410 per run). 8.3% binary completion. | 41.5% | Anthropic | Source | |
| Claude Sonnet 4.6 (medium thinking) New Author-run; medium thinking, standard tool calls, 500-step budget (~$1,550 per run). 9.3% binary completion (higher binary than max thinking). | 33.9% | Anthropic | Source | |
| MiniMax M3 New Author-run; reasoning enabled, standard tool calls, 500-step budget (~$259 per run). 4.6% binary completion. | 22.3% | MiniMax | Source | |
| Kimi 2.6 New Author-run; reasoning enabled, standard tool calls, 500-step budget (~$708 per run). 4.6% binary completion. | 22.1% | Moonshot AI | Source | |
| Qwen 3.7-Plus New Author-run; thinking mode, standard tool calls, 500-step budget (~$412 per run). 2.8% binary completion. | 21.5% | Alibaba | Source |
About this benchmark
OSWorld 2.0 evaluates computer-use agents on 108 long-horizon, end-to-end desktop workflows spanning everyday and professional tasks across self-hosted websites, office suites, files, and multi-application pipelines.
It is far harder than OSWorld 1.0: tasks take human users a median of about 1.6 hours, require an average of roughly 318 tool calls (vs about 30 in OSWorld 1.0), and 69.6% run longer than an hour, stressing cross-source reasoning, implicit-state inference, streaming interaction, and visual-spatial precision.
Scores here are not comparable to the OSWorld (1.0/Verified) leaderboard: the task set, step budgets, and metrics all differ, so use this page only for within-benchmark ranking.
Rows mix benchmark-author runs, vendor self-reports, and independent co-author runs (Snorkel AI, Simular); step budgets, releases, and graders differ, so read each row's note before comparing ranks.
Binary completion stays low (best tracked is 32.0%, Muse Spark 1.3) while partial scores pass 75%, so read both metrics; a partial score is not comparable to OSWorld 1.0/Verified pass rates.
Release mixing is the main trap: the 06.24 and 08.08 releases change tasks and grading, OpenAI's 08.08 numbers use the offline subset, and Anthropic states its Fable 5.1 run modified tasks and grading.
Example tasks
Three public tasks quoted from benchmark sources:
- "Please help me submit a reimbursement claim in the ExpenseFlow system." Citation: OSWorld 2.0 paper (Task 008)
- "Help me fill out this DS-2019 application for my J-1 student visa." Citation: OSWorld 2.0 paper
- "Go to the TravelHub booking page for Le Meurice and select the Deluxe Suite, stopping before the user enters personal information." Citation: OSWorld 2.0 paper (Task 052)
Methodology
- Tasks run in real VM desktop environments with execution-based validators that check final state against many fine-grained checkpoints (averaging 27.25 per task).
- OSWorld 2.0 scores two ways: binary completion (all checkpoints passed) and a partial score (fraction of checkpoints reached). We rank on the partial score — it differentiates systems far better than the low binary rates and matches how the GPT-5.6 result is reported — with each row's binary completion noted. Tracked scores use the default 500-step budget.
- Rows mix author runs on the official harness, vendor self-reports from system cards and launch posts, and independent runs from Snorkel AI and Simular (both benchmark co-authors). Reasoning effort, tool-call mode, and step budget are reported per row and materially affect scores.
- We track public results with source URLs and record who ran each result; independent reproductions not involving the benchmark authors are still rare.
Links
Related benchmarks
Compare this benchmark with related pages from the hub: