Define the unit

What counts as one accepted output: ________

Who checks it: ________

Measurement window: ________

Include the full boundary

Initial model calls: $____

Retry calls: $____

Retrieval and external tools: $____

Allocated infrastructure: $____

Review cost = review minutes / 60 x reviewer hourly rate: $____

Total cost = all the above: $____

Accepted outputs: ____

Cost per accepted output = total cost / accepted outputs: $____

If accepted outputs are zero, the unit cost is undefined. Report the spend and the failure instead of dividing by one.

Worked example, hypothetical

Initial calls cost $20, retries $5, retrieval $10, allocated infrastructure $15 and review $50. Total cost is $100. If 80 outputs are accepted, cost per accepted output is $1.25. These are invented teaching inputs, not provider prices or customer results.

Audit the assumptions

Record the price source and effective date. Distinguish measured bills from estimates. Record the allocation rule for shared infrastructure. Treat blocked, failed and abandoned requests explicitly. Avoid counting a retry both in the initial-call total and the retry line.

Compare cost alongside acceptance quality and latency. A cheaper API call can produce a more expensive useful result.