A score belongs to an experiment
A benchmark result describes an evaluated system under a particular set of conditions. It should identify the task set, benchmark release, model or agent version, harness, allowed tools, run budget, and scoring rule. Without those details, a percentage is difficult to interpret and easy to overgeneralize.
OSWorld V2 explicitly versions its evaluation components and warns against mixing release-controlled code, task assets, websites, and environment images. That is a concrete reason to retain the release identifier alongside each score rather than collapsing all results into a single OSWorld column.
Compare the same kind of measurement
An agent benchmark may measure task completion. A browser infrastructure evaluation may measure whether pages load or challenges are handled under a chosen test. The latter can matter to an agent, but it does not measure the agent’s reasoning or the correctness of its final business output.
Our editorial rule is to separate browser, desktop, and mobile results unless the benchmark methodology explicitly provides a comparable mixed evaluation. Do not average unrelated percentages into an overall intelligence score. A visual ranking can imply more than the underlying data supports even when every number is copied accurately.
Read the source relationship
A benchmark created by researchers can still be run by a vendor evaluating its own product. Distinguish ownership of the task set from ownership of the reported run. Academic, independent, vendor-run, and vendor-reported labels describe provenance; none automatically proves that two rows share the same conditions.
Look for missing information such as retry counts, excluded failures, human intervention, or cost accounting. Mark it unknown. If the harness differs, present the rows as contextual evidence and explain the difference instead of announcing a winner.
Use public results to design a pilot
A public benchmark can help identify candidates and failure modes to investigate. Your deployment decision still needs tasks that resemble your applications, identity model, deadlines, and consequential actions. A report-export workload and a desktop editing workload can stress different parts of the same system.
For your own pilot, preserve attempted-task counts, accepted outputs, retries, cost, latency, and intervention records. Freeze the task definition and environment before comparing candidates. If you change the harness, start a new series rather than silently extending the old chart. COMPUTER USE presents external evidence with attribution; it does not claim to have independently rerun the published evaluations.