VisualWebArena: methodology & results
Understand VisualWebArena Owner source snapshot 89f5af29305c: Execution-based visual web task success, evaluation environment and comparability limits.
- Release
- Owner source snapshot 89f5af29305c
- Environment
- browser
- Owner
- VisualWebArena research team
- Metric
- Execution-based visual web task success
What this evaluation measures
A controlled multimodal web-task benchmark. The full 910-task agent evaluation is distinct from the 233-task human-trajectory subset.
VisualWebArena evaluates browser tasks. Its reported metric is execution-based visual web task success across 910 tasks. Read the release-specific task definitions before treating this as a proxy for your own workflows.
Results belong to the exact system and evaluation setup named in the source. A vendor’s current API is not automatically the system used in a historical run. Hosted browser reachability, task completion and full desktop interaction are different measurements.
Pinned task and harness source
This definition pins the owner repository at 89f5af29305c3d1e9f97ce4421462060a70c9a03. It identifies a reproducible source snapshot, not a claim that every published historical score used this commit.
Observation and scoring boundaries
The owner provides screenshot/Set-of-Mark agent configuration and execution-based evaluators. Text-only agents, different action sets, and the 233-task human subset are different evaluation conditions.
Reproduction and unknowns
The supplied Classifieds script selects gpt-4-vision-preview, image_som observations and site-specific batch/reset settings. A full comparable result also needs an explicit maximum-step policy, the site images/reset state, all task IDs, exact evaluated model and matched scripts for the other sites. No historical numeric result is relabeled.
Attributed results
Can these results be compared?
No unified ranking is published. These records do not establish a controlled comparison across current tools.
- Match benchmark code, tasks, assets and environment versions.
- Check whether resets, retries, task exclusions and human interventions are counted.
- Compare the same metric: task success is different from reachability or cost per attempt.
- Re-run representative tasks using your own constraints before a purchase or deployment decision.
How to read computer-use benchmark evidence
Separate task success from your operating cost
Return to the Index
Sources & verification
Reviewed Sep 29, 2026. Architecture recommendations are editorial analysis; linked vendor documentation supports the underlying capability and safety facts.
- github.com/web-arena-x/visualwebarena/tree/89f5af29305c3d1e9f97ce4421462060a70c9a03
- raw.githubusercontent.com/web-arena-x/visualwebarena/89f5af29305c3d1e9f97ce4421462060a70c9a03/README.md
- github.com/web-arena-x/visualwebarena