Search COMPUTER USE

Search the evidence-backed index by name, vendor, or use case.

COMPUTER USE / BENCHMARK

VisualWebArena: methodology & results

Understand VisualWebArena Owner source snapshot 89f5af29305c: Execution-based visual web task success, evaluation environment and comparability limits.

Release
Owner source snapshot 89f5af29305c
Environment
browser
Owner
VisualWebArena research team
Metric
Execution-based visual web task success
Original methodology ↗

What this evaluation measures

A controlled multimodal web-task benchmark. The full 910-task agent evaluation is distinct from the 233-task human-trajectory subset.

Source snapshot identity is not a historical run revision. No score transfer across task subsets, environment images or benchmark families.

VisualWebArena evaluates browser tasks. Its reported metric is execution-based visual web task success across 910 tasks. Read the release-specific task definitions before treating this as a proxy for your own workflows.

Results belong to the exact system and evaluation setup named in the source. A vendor’s current API is not automatically the system used in a historical run. Hosted browser reachability, task completion and full desktop interaction are different measurements.

Pinned task and harness source

This definition pins the owner repository at 89f5af29305c3d1e9f97ce4421462060a70c9a03. It identifies a reproducible source snapshot, not a claim that every published historical score used this commit.

Observation and scoring boundaries

The owner provides screenshot/Set-of-Mark agent configuration and execution-based evaluators. Text-only agents, different action sets, and the 233-task human subset are different evaluation conditions.

Reproduction and unknowns

The supplied Classifieds script selects gpt-4-vision-preview, image_som observations and site-specific batch/reset settings. A full comparable result also needs an explicit maximum-step policy, the site images/reset state, all task IDs, exact evaluated model and matched scripts for the other sites. No historical numeric result is relabeled.

Attributed results

Numeric results are withheld: a version-matched result has not passed our evidence checks. This methodology guide remains useful for understanding scope and reproducibility.

Can these results be compared?

No unified ranking is published. These records do not establish a controlled comparison across current tools.

  • Match benchmark code, tasks, assets and environment versions.
  • Check whether resets, retries, task exclusions and human interventions are counted.
  • Compare the same metric: task success is different from reachability or cost per attempt.
  • Re-run representative tasks using your own constraints before a purchase or deployment decision.
RESEARCH NOTES

Sources & verification

Reviewed Sep 29, 2026. Architecture recommendations are editorial analysis; linked vendor documentation supports the underlying capability and safety facts.

  1. github.com/web-arena-x/visualwebarena/tree/89f5af29305c3d1e9f97ce4421462060a70c9a03
  2. raw.githubusercontent.com/web-arena-x/visualwebarena/89f5af29305c3d1e9f97ce4421462060a70c9a03/README.md
  3. github.com/web-arena-x/visualwebarena
How we verify evidence →