Search COMPUTER USE

Search the evidence-backed index by name, vendor, or use case.

COMPUTER USE / BENCHMARK

WebArena Verified: methodology & results

Understand WebArena Verified Owner source snapshot 6473f72db5dc: Deterministic task scoring from response and network trace, evaluation environment and comparability limits.

Release
Owner source snapshot 6473f72db5dc
Environment
browser
Owner
ServiceNow / WebArena Verified research team
Metric
Deterministic task scoring from response and network trace
Original methodology ↗

What this evaluation measures

An audited WebArena derivative with corrected tasks and deterministic evaluators. The full 812-task dataset and 258-task Hard subset are separate populations.

Source snapshot identity is not a historical run revision. No score transfer across task subsets, environment images or benchmark families.

WebArena Verified evaluates browser tasks. Its reported metric is deterministic task scoring from response and network trace across 812 tasks. Read the release-specific task definitions before treating this as a proxy for your own workflows.

Results belong to the exact system and evaluation setup named in the source. A vendor’s current API is not automatically the system used in a historical run. Hosted browser reachability, task completion and full desktop interaction are different measurements.

Versioned task identity

This definition pins owner commit 6473f72db5dcefc97b5725b59e734504edc28a21. The Hard subset manifest was retrieved at that revision and contains 258 unique task IDs; it is not the denominator for the full benchmark.

Evaluation artifacts

The owner evaluates structured agent responses and captured network traces, replacing LLM-as-judge and substring scoring with type-aware normalization and structural comparison. Offline evaluation does not remove the need for a correctly initialized run environment.

Environment and comparison boundaries

Owner documentation describes revised Docker environments and environment-control reset endpoints. Full and Hard results, original WebArena scores and older environment images must not share a comparability group. No verified per-run image/config bundle has been imported here.

Attributed results

Numeric results are withheld: a version-matched result has not passed our evidence checks. This methodology guide remains useful for understanding scope and reproducibility.

Can these results be compared?

No unified ranking is published. These records do not establish a controlled comparison across current tools.

  • Match benchmark code, tasks, assets and environment versions.
  • Check whether resets, retries, task exclusions and human interventions are counted.
  • Compare the same metric: task success is different from reachability or cost per attempt.
  • Re-run representative tasks using your own constraints before a purchase or deployment decision.
RESEARCH NOTES

Sources & verification

Reviewed Sep 29, 2026. Architecture recommendations are editorial analysis; linked vendor documentation supports the underlying capability and safety facts.

  1. github.com/ServiceNow/webarena-verified/tree/6473f72db5dcefc97b5725b59e734504edc28a21
  2. raw.githubusercontent.com/ServiceNow/webarena-verified/6473f72db5dcefc97b5725b59e734504edc28a21/README.md
  3. github.com/ServiceNow/webarena-verified
How we verify evidence →