An AI draft that appears in ten seconds can still take twenty minutes to make usable. If the manual task takes fifteen minutes, the fast generation did not save time. Measure the finished work rather than the first response.
Define useful output
Choose a repeatable task with a named owner. Write down what must be correct, what format the user needs and who can approve the result. For a meeting summary, acceptance might require decisions, responsible people and deadlines to match the approved notes. Fluency alone is not acceptance.
Pick representative cases before seeing the AI's results. Include difficult cases and ordinary failures. Avoid a comparison between a hard manual example and an unusually easy assisted example. Use approved, non-sensitive inputs for the initial pilot.
Measure the baseline
Record the manual time for each case, including the review normally needed before use. Note whether the output passes the acceptance rule. Keep the same start and end boundaries in the assisted comparison. If the manual process includes collecting inputs, the assisted process should include that collection too.
A small pilot can reveal obvious problems, but it does not establish performance for every employee or task. State the sample size and which cases were included. Treat differences in case difficulty as a limitation, not a hidden adjustment.
Count the assisted work
For each case, record preparation, waiting, human review and rework. Add them for total assisted minutes. Keep setup and training time as a separate cost that must eventually be recovered. Record failed or abandoned cases rather than silently excluding them.
Hypothetical example: manual work takes 30 minutes. Assisted preparation takes 3, waiting 1, review 12 and rework 8. Total assisted time is 24 minutes. The saving is 6 minutes. This example is arithmetic, not evidence that a real customer achieved that result.
If initial setup takes 120 minutes, a consistent six-minute saving needs 20 accepted cases to recover setup time. Subscription costs, monitoring and changes to the workflow still need their own accounting.
Review quality alongside time
Report the number attempted and the number accepted. Keep error categories, especially errors the reviewer nearly missed. A smaller time saving with dependable output can be more useful than a larger apparent saving that creates downstream correction work.
NIST's AI Risk Management Framework core provides a broader basis for defining oversight and evaluating system risks. This worksheet is a small operational measurement aid, not certification against that framework.
Make a decision you can revisit
At the end of the window, choose keep, change or stop. Explain the observed evidence. Save the prompt and model version, assign a reviewer and set the next review date. Re-measure after a material change.
Use the AI workflow worksheet, then read how to run a seven-day pilot. Subscribe to AI Operator Weekly for practical experiments with their limits stated.