OSWorld: methodology & results
Understand OSWorld 1.0 (historical; no normalized score imported): Task success rate (%), evaluation environment and comparability limits.
- Release
- 1.0 (historical; no normalized score imported)
- Environment
- desktop
- Owner
- XLang / OSWorld research team
- Metric
- Task success rate (%)
What this evaluation measures
Desktop task evaluation in real computer environments. Older reported scores require their original step budget, observation and harness settings.
OSWorld evaluates desktop tasks. Its reported metric is task success rate (%). Read the release-specific task definitions before treating this as a proxy for your own workflows.
Results belong to the exact system and evaluation setup named in the source. A vendor’s current API is not automatically the system used in a historical run. Hosted browser reachability, task completion and full desktop interaction are different measurements.
Attributed results
Can these results be compared?
No unified ranking is published. These records do not establish a controlled comparison across current tools.
- Match benchmark code, tasks, assets and environment versions.
- Check whether resets, retries, task exclusions and human interventions are counted.
- Compare the same metric: task success is different from reachability or cost per attempt.
- Re-run representative tasks using your own constraints before a purchase or deployment decision.
How to read computer-use benchmark evidence
Separate task success from your operating cost
Return to the Index
Sources & verification
Reviewed Sep 27, 2026. Architecture recommendations are editorial analysis; linked vendor documentation supports the underlying capability and safety facts.
How we verify evidence →