Search COMPUTER USE

Search the evidence-backed index by name, vendor, or use case.

COMPUTER USE / BENCHMARK

Browser Use Internal Bench Hard: methodology & results

Understand Browser Use Internal Bench Hard 106-task snapshot · updated 2026-08-01: Strict task success (%), evaluation environment and comparability limits.

Release
106-task snapshot · updated 2026-08-01
Environment
browser
Owner
Browser Use
Metric
Strict task success (%)
Original methodology ↗

What this evaluation measures

A vendor-run set of 106 live-web tasks. A task earns success only on an exact final answer or end state; cost is total recorded spend divided by solved tasks.

Private workload; model labels and harness details are reported by the vendor. This is not an independent ranking and does not transfer to current model versions.

Browser Use Internal Bench Hard evaluates browser tasks. Its reported metric is strict task success (%) across 106 tasks. Read the release-specific task definitions before treating this as a proxy for your own workflows.

Results belong to the exact system and evaluation setup named in the source. A vendor’s current API is not automatically the system used in a historical run. Hosted browser reachability, task completion and full desktop interaction are different measurements.

Attributed results

SYSTEM AS REPORTEDSCOREDATE / TYPESETUP & SOURCE
Browser Use82%

$0.17 / solved task

Aug 1, 2026
vendor run
Browser Use shared harness, 106 live-web tasks; release/commit not disclosed

Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task.

Original result ↗
Opus 562%

$3.40 / solved task

Aug 1, 2026
vendor run
Browser Use shared harness, 106 live-web tasks; release/commit not disclosed

Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task.

Original result ↗
Gemini 3.1 Pro59%

$2.20 / solved task

Aug 1, 2026
vendor run
Browser Use shared harness, 106 live-web tasks; release/commit not disclosed

Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task.

Original result ↗
Sonnet 559%

$1.55 / solved task

Aug 1, 2026
vendor run
Browser Use shared harness, 106 live-web tasks; release/commit not disclosed

Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task.

Original result ↗
GPT-5.652%

$1.10 / solved task

Aug 1, 2026
vendor run
Browser Use shared harness, 106 live-web tasks; release/commit not disclosed

Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task.

Original result ↗
Gemini 3.6 Flash46%

$0.62 / solved task

Aug 1, 2026
vendor run
Browser Use shared harness, 106 live-web tasks; release/commit not disclosed

Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task.

Original result ↗
GPT-537%

$0.44 / solved task

Aug 1, 2026
vendor run
Browser Use shared harness, 106 live-web tasks; release/commit not disclosed

Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task.

Original result ↗

Can these results be compared?

The listed records share the same recorded benchmark release, harness and source type. Other conditions may still differ; inspect their source notes before drawing a conclusion.

  • Match benchmark code, tasks, assets and environment versions.
  • Check whether resets, retries, task exclusions and human interventions are counted.
  • Compare the same metric: task success is different from reachability or cost per attempt.
  • Re-run representative tasks using your own constraints before a purchase or deployment decision.
RESEARCH NOTES

Sources & verification

Reviewed Sep 27, 2026. Architecture recommendations are editorial analysis; linked vendor documentation supports the underlying capability and safety facts.

  1. browser-use.com/benchmarks/agents
How we verify evidence →