Search COMPUTER USE

Search the evidence-backed index by name, vendor, or use case.

COMPUTER USE / BENCHMARK

ScreenSpot-Pro: methodology & results

Understand ScreenSpot-Pro arXiv:2504.07981v1 · English end-to-end baselines: GUI grounding accuracy (micro-average), evaluation environment and comparability limits.

Release
arXiv:2504.07981v1 · English end-to-end baselines
Environment
mixed
Owner
ScreenSpot-Pro research team
Metric
GUI grounding accuracy (micro-average)
Original methodology ↗

What this evaluation measures

Professional GUI grounding from fixed screenshots; selected results are historical English paper baselines.

This group excludes Chinese instructions, iterative methods, later submissions and workflow-success measurements.

ScreenSpot-Pro evaluates mixed tasks. Its reported metric is gui grounding accuracy (micro-average) across 1581 tasks. Read the release-specific task definitions before treating this as a proxy for your own workflows.

Results belong to the exact system and evaluation setup named in the source. A vendor’s current API is not automatically the system used in a historical run. Hosted browser reachability, task completion and full desktop interaction are different measurements.

Version and task boundary

Paper v1, published April 4, 2025: 1581 English screenshot tasks. Chinese experiments are separate.

Scoring and provenance

Predicted points must fall inside target boxes. OS-Atlas-7B and UGround figures agree between Tables 2 and 3. These are paper-reported results.

Limits and excluded records

Checkpoint hashes and run dates are undisclosed. GPT-4o remains withheld: tables report 0.8, narrative 0.9. No COMPUTER USE reproduction is claimed.

Attributed results

SYSTEM AS REPORTEDSCOREDATE / TYPESETUP & SOURCE
OS-Atlas-7B · original paper baseline18.9%Apr 4, 2025
academic
Paper v1, Tables 2–3; English grounding

Historical same-paper comparison only. Checkpoint hashes and run date undisclosed. Verified 2026-09-29; not a workflow-success score.

Original result ↗
UGround (7B) · original paper baseline16.5%Apr 4, 2025
academic
Paper v1, Tables 2–3; English grounding

Historical same-paper comparison only. Checkpoint hashes and run date undisclosed. Verified 2026-09-29; not a workflow-success score.

Original result ↗

Can these results be compared?

The listed records share the same recorded benchmark release, harness and source type. Other conditions may still differ; inspect their source notes before drawing a conclusion.

  • Match benchmark code, tasks, assets and environment versions.
  • Check whether resets, retries, task exclusions and human interventions are counted.
  • Compare the same metric: task success is different from reachability or cost per attempt.
  • Re-run representative tasks using your own constraints before a purchase or deployment decision.
RESEARCH NOTES

Sources & verification

Reviewed Sep 29, 2026. Architecture recommendations are editorial analysis; linked vendor documentation supports the underlying capability and safety facts.

  1. github.com/likaixin2000/ScreenSpot-Pro-GUI-Grounding
  2. arxiv.org/html/2504.07981v1
How we verify evidence →