A useful pilot produces a decision. It should show whether one defined workflow helps one group of people, under a stated review process. Seven days is a practical starting window, not a guarantee that the evidence is sufficient.
Day 1: choose one task
Choose recurring work with accessible inputs, a clear output and a person who knows what good work looks like. A draft internal FAQ may be a better starting point than an autonomous customer decision. Write the acceptance rule and the situations where a person must take over.
Name the owner. Define the data the tool may use, the people who can access it and what must stay out. Confirm the tool's approved handling arrangements before including business information.
Day 2: capture the manual baseline
Collect a small, representative set of cases. Record manual preparation, execution, review and correction time. Write down the errors and cases that cannot be completed. Do not select only cases that make the assisted workflow look good.
Save case identifiers rather than private content in the worksheet. The team should be able to match an identifier to approved evidence without copying customer records into a public example.
Day 3: write the assisted process
Choose one prompt and save its version. Define the output format, acceptable source material and review steps. Decide how a failed or incomplete response is handled. Keep external actions under a person's control during the pilot.
For example, the model can propose FAQ wording from approved documentation while a reviewer verifies each answer before publication. Record corrections; they are part of the work, not an exception to ignore.
Days 4 and 5: run and record
Use the same acceptance rule and time boundaries. Record preparation, waiting, review and rework for each attempted case. Keep failures. Avoid changing the prompt after every disappointing result without recording that change; each new version is a different intervention.
Pause if the workflow exposes restricted information, repeatedly produces unsupported answers or creates work that the reviewer cannot reliably check. A stopped pilot can be a useful result.
Day 6: compare the evidence
Compare total time and accepted outputs with the baseline. Include setup time, recurring tool cost and the practical burden on the reviewer. Distinguish observations from estimates. A sample of five easy cases is not evidence for every task the team performs.
Read measuring time savings after review before presenting a productivity claim. NIST's AI RMF core is useful background for assigning oversight and evaluating problems.
Day 7: decide and assign the next step
Choose keep, change or stop. Name the person responsible for the next action, the conditions that would reverse the decision and the next review date. If the sample is too small or cases are not comparable, write that conclusion instead of forcing a success claim.
Download the pilot worksheet. Subscribe to AI Operator Weekly for one focused workflow at a time.