WebArena Verified: methodology & results
Understand WebArena Verified Owner source snapshot 6473f72db5dc: Deterministic task scoring from response and network trace, evaluation environment and comparability limits.
- Release
- Owner source snapshot 6473f72db5dc
- Environment
- browser
- Owner
- ServiceNow / WebArena Verified research team
- Metric
- Deterministic task scoring from response and network trace
What this evaluation measures
An audited WebArena derivative with corrected tasks and deterministic evaluators. The full 812-task dataset and 258-task Hard subset are separate populations.
WebArena Verified evaluates browser tasks. Its reported metric is deterministic task scoring from response and network trace across 812 tasks. Read the release-specific task definitions before treating this as a proxy for your own workflows.
Results belong to the exact system and evaluation setup named in the source. A vendor’s current API is not automatically the system used in a historical run. Hosted browser reachability, task completion and full desktop interaction are different measurements.
Versioned task identity
This definition pins owner commit 6473f72db5dcefc97b5725b59e734504edc28a21. The Hard subset manifest was retrieved at that revision and contains 258 unique task IDs; it is not the denominator for the full benchmark.
Evaluation artifacts
The owner evaluates structured agent responses and captured network traces, replacing LLM-as-judge and substring scoring with type-aware normalization and structural comparison. Offline evaluation does not remove the need for a correctly initialized run environment.
Environment and comparison boundaries
Owner documentation describes revised Docker environments and environment-control reset endpoints. Full and Hard results, original WebArena scores and older environment images must not share a comparability group. No verified per-run image/config bundle has been imported here.
Attributed results
Can these results be compared?
No unified ranking is published. These records do not establish a controlled comparison across current tools.
- Match benchmark code, tasks, assets and environment versions.
- Check whether resets, retries, task exclusions and human interventions are counted.
- Compare the same metric: task success is different from reachability or cost per attempt.
- Re-run representative tasks using your own constraints before a purchase or deployment decision.
How to read computer-use benchmark evidence
Separate task success from your operating cost
Return to the Index
Sources & verification
Reviewed Sep 29, 2026. Architecture recommendations are editorial analysis; linked vendor documentation supports the underlying capability and safety facts.
- github.com/ServiceNow/webarena-verified/tree/6473f72db5dcefc97b5725b59e734504edc28a21
- raw.githubusercontent.com/ServiceNow/webarena-verified/6473f72db5dcefc97b5725b59e734504edc28a21/README.md
- github.com/ServiceNow/webarena-verified