Search COMPUTER USE

Search the evidence-backed index by name, vendor, or use case.

COMPUTER USE / GUIDE

Success rate vs cost per accepted task

A higher success rate and a lower per-run price describe different parts of the decision.

Define the accepted result first

Success rate is meaningful only after the task and acceptance rule are fixed. For a report export, acceptance might require the correct account, period, file format, and content. For a multi-step workflow, a partial result may be useful but still fail the agreed task. Keep partial completion visible without silently counting it as full success.

Our editorial cost definition is total measured cost for the evaluation divided by the number of accepted outcomes, when that denominator is nonzero. Include failed attempts and retries in the numerator. If no task is accepted, report that fact instead of presenting a finite cost per success.

Count the full resource path

A model charge is not necessarily the entire run cost. Browser or VM duration, proxy usage, agent service fees, storage, and review time may belong to separate accounts. Determine which components a published cost figure includes before comparing it with a figure from another source.

Keep currency and billing units explicit. A subscription can include allowances that change the marginal cost of one run without making the subscription free. A pilot should report its measured usage and the dated pricing assumptions separately so later readers can distinguish workload behavior from a changed vendor rate.

Compare at a shared operating budget

A system allowed many retries may improve completion while increasing cost and latency. Another may stop early and return an exception for review. Neither behavior can be evaluated through success rate alone. Freeze the retry, time, and step limits when comparing candidates, then inspect the tradeoff under a deliberate second budget if useful.

Human review also changes the result. Record whether intervention is counted as failure, assisted completion, or an allowed part of the workflow. A high assisted-success rate can be valuable, but it should not be advertised as unattended performance. Review minutes can matter more than small differences in model spend.

Use a decision frontier, not one score

For your workload, compare accepted output, total cost, elapsed time, and severity of failures. Eliminate options that violate a hard constraint before optimizing the remainder. A cheap workflow that occasionally targets the wrong customer is not rescued by a favorable average cost.

Published agent benchmarks can provide useful context, but they do not establish your production economics. The external results in this atlas retain source attribution; no hypothetical calculation here is presented as a measured vendor outcome. Use the site’s dated pricing records for planning, then replace assumptions with a controlled pilot and keep uncertainty visible.

RESEARCH NOTES

Sources & verification

Reviewed Sep 27, 2026. Architecture recommendations are editorial analysis; linked vendor documentation supports the underlying capability and safety facts.

  1. browser-use.com/benchmarks/agents
  2. developers.openai.com/api/docs/guides/tools-computer-use
  3. github.com/xlang-ai/OSWorld-V2
How we verify evidence →

Put the guidance to work