Search COMPUTER USE

Search the evidence-backed index by name, vendor, or use case.

COMPUTER USE / BENCHMARK

WebArena: methodology & results

Understand WebArena Task annotations v0.2.0 · 2023-10-24: Execution-based task success, evaluation environment and comparability limits.

Release
Task annotations v0.2.0 · 2023-10-24
Environment
browser
Owner
WebArena research team
Metric
Execution-based task success
Original methodology ↗

What this evaluation measures

Owner describes 812 controlled-site tasks; v0.2.0 updates annotations. Reproducibility requires self-hosted/reset environments rather than mutable demo sites.

WebArenaVerified and VisualWebArena are different suites. Live demo websites are not resettable benchmark environments; task annotation version alone is insufficient for a comparable result.

WebArena evaluates browser tasks. Its reported metric is execution-based task success across 812 tasks. Read the release-specific task definitions before treating this as a proxy for your own workflows.

Results belong to the exact system and evaluation setup named in the source. A vendor’s current API is not automatically the system used in a historical run. Hosted browser reachability, task completion and full desktop interaction are different measurements.

Owner methodology and measurement boundary

Owner describes 812 controlled-site tasks; v0.2.0 updates annotations. Reproducibility requires self-hosted/reset environments rather than mutable demo sites.

Reproducibility requirements

Pin code, task JSON, environment images and scoring configuration.

Version history and interpretation

v0.2.0 updated task annotations in October 2023. WebArenaVerified changes evaluation and includes a Hard subset; its scores must not be placed in this original suite's leaderboard.

Attributed results

Numeric results are withheld: a version-matched result has not passed our evidence checks. This methodology guide remains useful for understanding scope and reproducibility.

Can these results be compared?

No unified ranking is published. These records do not establish a controlled comparison across current tools.

  • Match benchmark code, tasks, assets and environment versions.
  • Check whether resets, retries, task exclusions and human interventions are counted.
  • Compare the same metric: task success is different from reachability or cost per attempt.
  • Re-run representative tasks using your own constraints before a purchase or deployment decision.
RESEARCH NOTES

Sources & verification

Reviewed Sep 28, 2026. Architecture recommendations are editorial analysis; linked vendor documentation supports the underlying capability and safety facts.

  1. github.com/web-arena-x/webarena
How we verify evidence →