Search COMPUTER USE

Search the evidence-backed index by name, vendor, or use case.

COMPUTER USE / HUB

Benchmark Atlas

Read results in context: task, version, harness, date and source ownership.

An atlas of external evidence, not an overall leaderboard. Scores keep their original system names, versions and source attribution.

Reported results

BrowserBench / Browser Use run · 2026-03-21; upstream revision undisclosedvendor run

Browser Use Cloud

84.8%

Anti-bot access success (%)

Reported Mar 21, 2026 · Halluminate BrowserBench run by Browser Use; upstream revision not disclosed

Provider label exactly as reported. Metric measures anti-bot access. COMPUTER USE did not run this evaluation.

Read the original result ↗Methodology →
BrowserBench / Browser Use run · 2026-03-21; upstream revision undisclosedvendor run

Hyperbrowser

76.4%

Anti-bot access success (%)

Reported Mar 21, 2026 · Halluminate BrowserBench run by Browser Use; upstream revision not disclosed

Provider label exactly as reported. Metric measures anti-bot access. COMPUTER USE did not run this evaluation.

Read the original result ↗Methodology →
BrowserBench / Browser Use run · 2026-03-21; upstream revision undisclosedvendor run

Anchor

76%

Anti-bot access success (%)

Reported Mar 21, 2026 · Halluminate BrowserBench run by Browser Use; upstream revision not disclosed

Provider label exactly as reported. Metric measures anti-bot access. COMPUTER USE did not run this evaluation.

Read the original result ↗Methodology →
BrowserBench / Browser Use run · 2026-03-21; upstream revision undisclosedvendor run

Steel

73.3%

Anti-bot access success (%)

Reported Mar 21, 2026 · Halluminate BrowserBench run by Browser Use; upstream revision not disclosed

Provider label exactly as reported. Metric measures anti-bot access. COMPUTER USE did not run this evaluation.

Read the original result ↗Methodology →
BrowserBench / Browser Use run · 2026-03-21; upstream revision undisclosedvendor run

Browserbase

70.3%

Anti-bot access success (%)

Reported Mar 21, 2026 · Halluminate BrowserBench run by Browser Use; upstream revision not disclosed

Provider label exactly as reported. Metric measures anti-bot access. COMPUTER USE did not run this evaluation.

Read the original result ↗Methodology →
Browser Use Internal Bench Hard / 106-task snapshot · updated 2026-08-01vendor run

Browser Use

82%

Strict task success (%)

$0.17 / solved task (source-reported)

Reported Aug 1, 2026 · Browser Use shared harness, 106 live-web tasks; release/commit not disclosed

Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task.

Read the original result ↗Methodology →
Browser Use Internal Bench Hard / 106-task snapshot · updated 2026-08-01vendor run

Opus 5

62%

Strict task success (%)

$3.40 / solved task (source-reported)

Reported Aug 1, 2026 · Browser Use shared harness, 106 live-web tasks; release/commit not disclosed

Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task.

Read the original result ↗Methodology →
Browser Use Internal Bench Hard / 106-task snapshot · updated 2026-08-01vendor run

Gemini 3.1 Pro

59%

Strict task success (%)

$2.20 / solved task (source-reported)

Reported Aug 1, 2026 · Browser Use shared harness, 106 live-web tasks; release/commit not disclosed

Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task.

Read the original result ↗Methodology →
Browser Use Internal Bench Hard / 106-task snapshot · updated 2026-08-01vendor run

Sonnet 5

59%

Strict task success (%)

$1.55 / solved task (source-reported)

Reported Aug 1, 2026 · Browser Use shared harness, 106 live-web tasks; release/commit not disclosed

Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task.

Read the original result ↗Methodology →
Browser Use Internal Bench Hard / 106-task snapshot · updated 2026-08-01vendor run

GPT-5.6

52%

Strict task success (%)

$1.10 / solved task (source-reported)

Reported Aug 1, 2026 · Browser Use shared harness, 106 live-web tasks; release/commit not disclosed

Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task.

Read the original result ↗Methodology →
Browser Use Internal Bench Hard / 106-task snapshot · updated 2026-08-01vendor run

Gemini 3.6 Flash

46%

Strict task success (%)

$0.62 / solved task (source-reported)

Reported Aug 1, 2026 · Browser Use shared harness, 106 live-web tasks; release/commit not disclosed

Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task.

Read the original result ↗Methodology →
Browser Use Internal Bench Hard / 106-task snapshot · updated 2026-08-01vendor run

GPT-5

37%

Strict task success (%)

$0.44 / solved task (source-reported)

Reported Aug 1, 2026 · Browser Use shared harness, 106 live-web tasks; release/commit not disclosed

Exact model label from source; no precise model API revision disclosed. Historical score is not attached to a current category model profile. Cost in USD per solved task.

Read the original result ↗Methodology →

Reading the evidence

Explore benchmark methodology