ScreenSpot-Pro: methodology & results
Understand ScreenSpot-Pro arXiv:2504.07981v1 · English end-to-end baselines: GUI grounding accuracy (micro-average), evaluation environment and comparability limits.
- Release
- arXiv:2504.07981v1 · English end-to-end baselines
- Environment
- mixed
- Owner
- ScreenSpot-Pro research team
- Metric
- GUI grounding accuracy (micro-average)
What this evaluation measures
Professional GUI grounding from fixed screenshots; selected results are historical English paper baselines.
ScreenSpot-Pro evaluates mixed tasks. Its reported metric is gui grounding accuracy (micro-average) across 1581 tasks. Read the release-specific task definitions before treating this as a proxy for your own workflows.
Results belong to the exact system and evaluation setup named in the source. A vendor’s current API is not automatically the system used in a historical run. Hosted browser reachability, task completion and full desktop interaction are different measurements.
Version and task boundary
Paper v1, published April 4, 2025: 1581 English screenshot tasks. Chinese experiments are separate.
Scoring and provenance
Predicted points must fall inside target boxes. OS-Atlas-7B and UGround figures agree between Tables 2 and 3. These are paper-reported results.
Limits and excluded records
Checkpoint hashes and run dates are undisclosed. GPT-4o remains withheld: tables report 0.8, narrative 0.9. No COMPUTER USE reproduction is claimed.
Attributed results
| SYSTEM AS REPORTED | SCORE | DATE / TYPE | SETUP & SOURCE |
|---|---|---|---|
| OS-Atlas-7B · original paper baseline | 18.9% | Apr 4, 2025 academic | Paper v1, Tables 2–3; English grounding Historical same-paper comparison only. Checkpoint hashes and run date undisclosed. Verified 2026-09-29; not a workflow-success score. Original result ↗ |
| UGround (7B) · original paper baseline | 16.5% | Apr 4, 2025 academic | Paper v1, Tables 2–3; English grounding Historical same-paper comparison only. Checkpoint hashes and run date undisclosed. Verified 2026-09-29; not a workflow-success score. Original result ↗ |
Can these results be compared?
The listed records share the same recorded benchmark release, harness and source type. Other conditions may still differ; inspect their source notes before drawing a conclusion.
- Match benchmark code, tasks, assets and environment versions.
- Check whether resets, retries, task exclusions and human interventions are counted.
- Compare the same metric: task success is different from reachability or cost per attempt.
- Re-run representative tasks using your own constraints before a purchase or deployment decision.
How to read computer-use benchmark evidence
Separate task success from your operating cost
Return to the Index
Sources & verification
Reviewed Sep 29, 2026. Architecture recommendations are editorial analysis; linked vendor documentation supports the underlying capability and safety facts.
How we verify evidence →