A user sees one task. Your system may make several model or tool calls before producing an answer. If the cost report counts only the final response, retries disappear from the budget. If it counts attempts and then adds the same retries again, the budget is overstated. Start with a clear event model.

Give tasks and attempts different identifiers

Assign an internal task ID to the requested outcome. Give every external attempt its own ID. Record the provider request ID when available, attempt number, start and end time, state and usage evidence. Do not place prompts, credentials or customer text into a public log.

For a hypothetical example, task T1 has attempts A1, A2 and A3. A1 times out, A2 fails review and A3 is accepted. There is one user task, three attempts and one accepted output. The timeout does not establish that the provider performed no work or incurred no cost.

Keep the failure reason

Distinguish provider rejection, network uncertainty, invalid output, review failure and a user cancellation. They need different responses. An uncertain timeout may need reconciliation before another external action; a formatting error may need a bounded correction attempt.

Save the retry policy and limit before running the comparison. Avoid endless retries until a case passes. A bounded failure is easier to inspect than a hidden series of repeated calls.

Reconcile usage evidence

Match returned usage and provider records to attempts where possible. Keep missing usage marked unknown or estimated. Record the source and date of any price estimate. Provider billing can arrive later than application events, so record when the reconciliation was performed.

Choose one accounting method: sum all attempts, or show initial and retry costs as mutually exclusive components. Never add a retry subtotal on top of an all-attempt total. Include tools and retrieval calls within the boundary you defined.

Review cost and useful outcomes together

Group attempts by task and failure category. Ask which category causes the most repeated work and which tasks consume review time without producing an acceptable result. Compare the same task mix across versions.

In a hypothetical sample of 100 tasks, 20 tasks need two attempts and 80 need one. That is 120 attempts, not 100. Whether those attempts cost the same depends on measured usage and request classes. The example cannot supply a real bill without those records.

The FinOps Foundation's AI tools and services considerations discusses use-case costs in relation to the outcome being produced. Read budgeting by useful output for the full cost boundary.

Change one thing, then measure again

Pick the dominant failure category. Change one validation, prompt or retry rule, save the version and compare the next representative sample. Report the effect on accepted outputs, cost and latency together.

Download the AI request cost worksheet. Subscribe to Cloud & AI Engineering Brief for evidence-based operating notes.