Token price is one input to an AI budget. The engineering question is what it costs to produce an output someone can actually use. An inexpensive response that fails review can be more costly than a larger response that needs little correction.

Define the useful unit

Write an acceptance rule before measuring. Your unit might be an approved support draft or a reviewed research summary. An API request, a user task and an accepted output are different denominators. State which one the budget uses.

The FinOps Foundation's unit economics capability discusses connecting technology costs to meaningful units. Apply that idea to the outcome your team needs, rather than treating token volume as the outcome.

Draw the cost boundary

Include initial model calls, retry calls, retrieval, external tools, allocated infrastructure and human review. State whether setup, evaluation and monitoring are inside the period. Keep the boundary consistent when comparing configurations.

Record actual billed usage when available. If a cost is estimated, label it estimated and save the price source and effective date. Do not treat a dashboard estimate as a reconciled invoice. Account for cached input or different request classes only when the records identify them.

Calculate the unit cost

Total cost equals the included cost components. Cost per accepted output equals total cost divided by accepted outputs. If no outputs are accepted, the unit cost is undefined; report the spend and the failure.

A hypothetical monthly example has $20 of initial calls, $5 of retries, $10 of retrieval, $15 of allocated infrastructure and $50 of review time. Total cost is $100. With 80 accepted outputs, unit cost is $1.25. These numbers are invented teaching inputs, not provider prices or customer results.

If only 50 outputs are accepted at the same $100 cost, the useful-output cost becomes $2. The change came from acceptance, not the model's token price. Track both.

Allocate shared costs openly

A shared server does not usually provide a perfect bill for each request. Choose an allocation rule, such as measured usage or a documented share, and record its limitations. Do not silently assign all fixed costs to one test run or ignore them in another.

Review time can dominate a small pilot. Convert minutes to cost using an explicit hourly assumption. If review is already included in a separate staff-cost allocation, do not add it again.

Compare configurations fairly

Use the same representative cases and acceptance rule. Compare cost with latency, failures and review effort. Keep rejected and abandoned work visible. A cheaper configuration that shifts correction work to another team may not reduce the full cost.

Use the cost-per-useful-request worksheet and read how retries change the AI bill. Subscribe to Cloud & AI Engineering Brief for practical operating methods.