Benchmark

OSWorld 2.0

OSWorld 2.0 is a benchmark of 108 long-horizon, real-world computer-use workflows spanning everyday and professional tasks. Each task is an end-to-end…

Modality
multimodal
Categories
multimodal, general, agents, vision
Openness
unknown
Source
llm_stats
Reported scores
5

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Opus 5Anthropic0.7062026-07-24evidence
Gemini 3.7 FlashGoogle0.4792026-08-13evidence
GPT-5.6 LunaOpenAI0.4562026-07-09evidence
GPT-5.6 SolOpenAI0.6262026-07-09evidence
GPT-5.6 TerraOpenAI0.5022026-07-09evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.