the inference company

The Inference Company / Cost per successful task

Cost per successful task

Cost per successful task is the total you spend running an agent on a set of tasks, divided by the number of those tasks that succeeded. Failed attempts still cost money, so their spend is carried by the tasks that worked.

The formula

cost per successful task = spend on tasks with an outcome ÷ tasks that succeeded

Two rules keep the number honest:

  • Failures count on the spend side. A task that failed still used tokens and GPU time. Its cost is added to the total, and only successes go in the denominator.
  • Tasks without an outcome are left out. If nobody recorded whether a task worked, it is excluded from both sides and reported separately as unknown.

A worked example

A coding agent runs three tasks. Success means the tests pass.

TaskResultCost
Fix flaky cache testTests pass$0.41
Add retry to clientTests pass$0.58
Upgrade auth libraryTests fail$0.77

Example numbers, for illustration.

Total spend is $1.76 and two tasks succeeded, so the cost per successful task is $0.88. The average cost per task is $0.59, but that number counts the failed upgrade as if it were a result. You paid $0.88 for each thing that actually worked.

Why price per token is not enough

Token prices describe the inputs. An agent's real cost depends on how many turns it takes, how much context it resends each turn, how often it retries, and how often it fails outright. A setup with cheaper tokens can still cost more per working result if it needs more turns or succeeds less often.

Cost per successful task puts all of that into one number you can compare across models, providers and agent versions.

What counts as success

Use the signal your agent already has:

  • Coding agents: the test suite passes, or the pull request is merged.
  • Support agents: the ticket is resolved without being handed to a person.
  • Research agents: the report is accepted.

Define it per agent. Comparing agents is only fair when each one is judged by the outcome it exists to produce.

How to measure it

You need two things on every task: its spend, and its outcome.

1. Tag each request with a task ID

Send your own ID (a ticket number, a pull request) with every model call that belongs to the task, so spend adds up per task.

# any OpenAI or Anthropic client
headers = {"x-task-id": "ticket-4821"}

2. Report the outcome when the task ends

POST /v1/tasks/ticket-4821/finish
{"status": "completed", "success": true, "success_source": "ticket resolved"}

On The Inference Company, that is all the setup there is: spend is recorded on every request, and the dashboard shows cost per successful task by agent and over time. With any other provider, the same two pieces of data give you the same number.

What brings it down

  • Not paying twice for the same context. Agents resend most of their conversation every turn. Keeping it ready between turns, instead of rebuilding it, removes repeated work.
  • Running blocked work first. A call carrying a tool result is holding up a task in progress; serving it ahead of background calls shortens the task.
  • Changing serving only when results hold. A change that saves money but makes tasks fail more often raises cost per successful task. Test changes on replayed agent sessions before shipping them.

Common questions

Is cost per successful task the same as average cost per task?

No. Average cost per task divides spend by every attempt, so failures make the number look cheaper than a working result really is. Cost per successful task divides by successes only, so failed attempts raise it.

Do tasks without an outcome count?

No. A task with no recorded success or failure is left out of both the spend and the count, and reported separately as unknown, so it can't make the number look better or worse than it is.

Why not just compare price per token?

Because a cheaper token can still mean a more expensive result. If a model or setup needs more turns, more retries or fails more often, the price per token falls while the cost of each working result rises.

What counts as a successful task?

Whatever signal your agent already has: tests that pass for a coding agent, a resolved ticket for a support agent, an accepted report for a research agent. You define it per agent.

Want to know yours?

We're onboarding a small group of teams running agents on open models. Book a call and we'll show you how we'd measure cost per successful task against your current provider.