Search COMPUTER USE

Search the evidence-backed index by name, vendor, or use case.

COMPUTER USE / GUIDE

How to compare benchmark results

Use a compatibility checklist before putting two published scores in the same ranking.

Start with the experiment identity

Write down the benchmark name and release, task population, metric definition, evaluated system, and run date for each result. Then add the harness, environment, allowed tools, step or time budget, and retry policy. A shared benchmark name is only the beginning of comparability, especially when the task set changes over time.

OSWorld V2 explicitly coordinates versioned evaluation components. Browser Use separately publishes agent and browser-infrastructure evaluations. These are practical reasons to retain both version and measurement type in a data model rather than forcing every percentage into one “success” field.

Check the numerator and denominator

Ask what counts as a success and which attempts are included. A task-level metric may require all subtasks to pass, while another result may score individual actions. A subset chosen for difficulty or availability can differ from the full task population. An excluded failed run changes the denominator even if the displayed percentage is calculated correctly.

Our editorial rule is to show unknown when a report does not disclose a required detail. Do not fill in an assumed default and call the rows comparable. A transparent evidence table can be useful even when the conclusion is that the available information does not support a ranking.

Separate comparability from credibility

An independent source can publish a result with an incompatible harness. A vendor can publish a well-documented result that is comparable within a defined series. Source relationship matters, but it answers a different question from whether the experimental conditions match.

Read who selected the tasks, ran the system, scored outputs, and paid for or published the evaluation. Then inspect artifacts and failure accounting where available. Avoid converting this review into an invented confidence percentage; explain the specific evidence that is present and the conditions still missing.

Choose the right presentation

When the conditions match, a side-by-side table can make the tradeoffs clear. When they do not, separate the rows into labeled groups and state the mismatch near the score. Do not rely on a distant footnote to undo the impression created by a single ordered chart.

For a buying decision, use the public evidence to design a task-specific pilot with frozen conditions. Retain cost, latency, intervention, and failure notes alongside completion. A result can support a narrow claim about a defined evaluation without supporting a universal winner across browser, desktop, mobile, and business workflows.

RESEARCH NOTES

Sources & verification

Reviewed Sep 27, 2026. Architecture recommendations are editorial analysis; linked vendor documentation supports the underlying capability and safety facts.

  1. github.com/xlang-ai/OSWorld-V2
  2. browser-use.com/benchmarks/agents
  3. browser-use.com/benchmarks/browsers
How we verify evidence →

Put the guidance to work