About this benchmark

OSWorld 2.0 evaluates computer-use agents on 108 long-horizon, end-to-end desktop workflows spanning everyday and professional tasks across self-hosted websites, office suites, files, and multi-application pipelines.

It is far harder than OSWorld 1.0: tasks take human users a median of about 1.6 hours, require an average of roughly 318 tool calls (vs about 30 in OSWorld 1.0), and 69.6% run longer than an hour, stressing cross-source reasoning, implicit-state inference, streaming interaction, and visual-spatial precision.

Scores here are not comparable to the OSWorld (1.0/Verified) leaderboard: the task set, step budgets, and metrics all differ, so use this page only for within-benchmark ranking.

Rows mix benchmark-author runs, vendor self-reports, and independent co-author runs (Snorkel AI, Simular); step budgets, releases, and graders differ, so read each row's note before comparing ranks.

Binary completion stays low (best tracked is 32.0%, Muse Spark 1.3) while partial scores pass 75%, so read both metrics; a partial score is not comparable to OSWorld 1.0/Verified pass rates.

Release mixing is the main trap: the 06.24 and 08.08 releases change tasks and grading, OpenAI's 08.08 numbers use the offline subset, and Anthropic states its Fable 5.1 run modified tasks and grading.

Example tasks

Three public tasks quoted from benchmark sources:

Methodology

  • Tasks run in real VM desktop environments with execution-based validators that check final state against many fine-grained checkpoints (averaging 27.25 per task).
  • OSWorld 2.0 scores two ways: binary completion (all checkpoints passed) and a partial score (fraction of checkpoints reached). We rank on the partial score — it differentiates systems far better than the low binary rates and matches how the GPT-5.6 result is reported — with each row's binary completion noted. Tracked scores use the default 500-step budget.
  • Rows mix author runs on the official harness, vendor self-reports from system cards and launch posts, and independent runs from Snorkel AI and Simular (both benchmark co-authors). Reasoning effort, tool-call mode, and step budget are reported per row and materially affect scores.
  • We track public results with source URLs and record who ran each result; independent reproductions not involving the benchmark authors are still rare.

Related benchmarks

Compare this benchmark with related pages from the hub:

Back to benchmark hub

Frequently asked questions

Which system is currently best on OSWorld 2.0? +
Claude Fable 5.1 is the system/agent setup currently leading with a tracked score of 77.9%. This ranking reflects submitted system setups (model plus tools and policy), not just a base model. Based on our latest tracked results, last updated Sep 30, 2026.
What should I read into a OSWorld 2.0 score? +
OSWorld 2.0 scores are most useful for within-benchmark ranking. Read the Notes column to understand setup context, and use the methodology section before making procurement or architecture decisions.
Are these independently verified? +
Not always. Some rows are independently benchmarked and some are team-reported. Use each source link and notes field to verify evidence level before drawing strong conclusions.