The dataset owner is not always the runner
An academic benchmark can be used in a vendor’s self-evaluation. An independent evaluator can test a commercial model with its own harness. A vendor can publish a result reported by another organization. These relationships should remain distinct instead of being compressed into a single “official benchmark” badge.
In this atlas, source labels describe the reporting relationship: academic, independent, vendor-run, or vendor-reported. They are provenance labels, not a numeric credibility score. Browser Use’s own published evaluations belong in the vendor-run context when Browser Use conducted the run, even if an underlying task set came from elsewhere.
Ask four ownership questions
Who chose or created the tasks? Who configured and ran the system? Who evaluated the outputs? Who published or funded the result? A single organization may hold every role, or the roles may be split. Record what the source actually discloses and leave the rest unknown.
Our editorial review then looks for task definitions, version identifiers, run budgets, failure accounting, and inspectable artifacts. These details make a result more useful to a reader. A label such as independent cannot compensate for a missing metric definition, and a vendor relationship does not by itself invalidate disclosed experimental data.
Look for selection effects
A published subset may emphasize tasks that are difficult, commercially relevant, or compatible with a particular environment. That can be a legitimate experiment, but the selection needs to be visible. Likewise, an evaluation may use a specific model configuration, custom prompts, or extra tools that differ from a default customer setup.
Ask whether the result describes a product as shipped, a research configuration, or an optimized demonstration. Keep those records separate. Do not turn a carefully tuned benchmark result into a promise that a new account will achieve the same performance on unrelated portals.
Use attribution without inventing a verdict
Display the runner relationship next to each score, with a direct source and date. When reports disagree, first compare versions and conditions rather than choosing the more flattering number. Preserve historical records so a later change does not make an earlier claim appear to have used the newer setup.
For procurement, treat both vendor and independent evidence as inputs to your own bounded pilot. Require an acceptance rule that reflects the actual business task. COMPUTER USE aggregates public evidence and explains its limits; it does not call a vendor’s run an independent test or imply that it has reproduced results it has only inspected.