Browser Use Internal Bench Hard: methodology & results
Understand Browser Use Internal Bench Hard 106-task snapshot · updated 2026-08-01: Strict task success (%), evaluation environment and comparability limits.
- Release
- 106-task snapshot · updated 2026-08-01
- Environment
- browser
- Owner
- Browser Use
- Metric
- Strict task success (%)
What this evaluation measures
A vendor-run set of 106 live-web tasks. A task earns success only on an exact final answer or end state; cost is total recorded spend divided by solved tasks.
Browser Use Internal Bench Hard evaluates browser tasks. Its reported metric is strict task success (%) across 106 tasks. Read the release-specific task definitions before treating this as a proxy for your own workflows.
Results belong to the exact system and evaluation setup named in the source. A vendor’s current API is not automatically the system used in a historical run. Hosted browser reachability, task completion and full desktop interaction are different measurements.
Attributed results
| SYSTEM AS REPORTED | SCORE | DATE / TYPE | SETUP & SOURCE |
|---|---|---|---|
| Browser Use | 82% $0.17 / solved task | Aug 1, 2026 vendor run | Browser Use shared harness, 106 live-web tasks; release/commit not disclosed Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task. Original result ↗ |
| Opus 5 | 62% $3.40 / solved task | Aug 1, 2026 vendor run | Browser Use shared harness, 106 live-web tasks; release/commit not disclosed Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task. Original result ↗ |
| Gemini 3.1 Pro | 59% $2.20 / solved task | Aug 1, 2026 vendor run | Browser Use shared harness, 106 live-web tasks; release/commit not disclosed Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task. Original result ↗ |
| Sonnet 5 | 59% $1.55 / solved task | Aug 1, 2026 vendor run | Browser Use shared harness, 106 live-web tasks; release/commit not disclosed Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task. Original result ↗ |
| GPT-5.6 | 52% $1.10 / solved task | Aug 1, 2026 vendor run | Browser Use shared harness, 106 live-web tasks; release/commit not disclosed Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task. Original result ↗ |
| Gemini 3.6 Flash | 46% $0.62 / solved task | Aug 1, 2026 vendor run | Browser Use shared harness, 106 live-web tasks; release/commit not disclosed Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task. Original result ↗ |
| GPT-5 | 37% $0.44 / solved task | Aug 1, 2026 vendor run | Browser Use shared harness, 106 live-web tasks; release/commit not disclosed Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task. Original result ↗ |
Can these results be compared?
The listed records share the same recorded benchmark release, harness and source type. Other conditions may still differ; inspect their source notes before drawing a conclusion.
- Match benchmark code, tasks, assets and environment versions.
- Check whether resets, retries, task exclusions and human interventions are counted.
- Compare the same metric: task success is different from reachability or cost per attempt.
- Re-run representative tasks using your own constraints before a purchase or deployment decision.
How to read computer-use benchmark evidence
Separate task success from your operating cost
Return to the Index
Sources & verification
Reviewed Sep 27, 2026. Architecture recommendations are editorial analysis; linked vendor documentation supports the underlying capability and safety facts.
How we verify evidence →