Search COMPUTER USE

Search the evidence-backed index by name, vendor, or use case.

COMPUTER USE / BENCHMARK

OSWorld: methodology & results

Understand OSWorld 1.0 (historical; no normalized score imported): Task success rate (%), evaluation environment and comparability limits.

Release
1.0 (historical; no normalized score imported)
Environment
desktop
Owner
XLang / OSWorld research team
Metric
Task success rate (%)
Original methodology ↗

What this evaluation measures

Desktop task evaluation in real computer environments. Older reported scores require their original step budget, observation and harness settings.

OSWorld 1.0 results are not comparable to OSWorld V2 scores; this atlas has not normalized a numeric 1.0 result.

OSWorld evaluates desktop tasks. Its reported metric is task success rate (%). Read the release-specific task definitions before treating this as a proxy for your own workflows.

Results belong to the exact system and evaluation setup named in the source. A vendor’s current API is not automatically the system used in a historical run. Hosted browser reachability, task completion and full desktop interaction are different measurements.

Attributed results

Numeric results are withheld: a version-matched result has not passed our evidence checks. This methodology guide remains useful for understanding scope and reproducibility.

Can these results be compared?

No unified ranking is published. These records do not establish a controlled comparison across current tools.

  • Match benchmark code, tasks, assets and environment versions.
  • Check whether resets, retries, task exclusions and human interventions are counted.
  • Compare the same metric: task success is different from reachability or cost per attempt.
  • Re-run representative tasks using your own constraints before a purchase or deployment decision.
RESEARCH NOTES

Sources & verification

Reviewed Sep 27, 2026. Architecture recommendations are editorial analysis; linked vendor documentation supports the underlying capability and safety facts.

  1. github.com/xlang-ai/OSWorld
How we verify evidence →