All Insights
AI at work5 min read

Measure an AI pilot by accepted work and total effort

Count preparation, review, corrections, fallback and maintenance. Use a worked example and a baseline worksheet to decide whether an AI pilot saves time.

Illustrative human effort for 20 completed tasks: 240 minutes manually versus 180 minutes of assisted task work plus 20 minutes of maintenance, saving 40 minutes.
Fictional comparison for 20 accepted tasks. Initial setup is tracked separately; these are not measured results. Open full-size diagram

Measure an AI pilot by the total work needed to produce an acceptable result. Include preparation, review, correction, manual fallback and ongoing maintenance. Generation speed alone cannot tell you whether the team saved time.

Define “acceptable” before comparing runs. A job summary that omits a return visit is not equivalent to a complete manual summary. An invoice check that invents a charge does not become successful because it arrived quickly.

NIST's voluntary AI framework calls for evaluation in the relevant context and continued monitoring. The worksheet below is a proposed way to apply that discipline to one business task; it is not a NIST-mandated benchmark or a proven DATUM result. NIST AI RMF Core.

Compare the same finished task

Pick a bounded result that the responsible person can judge. For example: prepare a job-readiness note that identifies missing materials and unresolved approvals, with references to the source records.

List the conditions for accepting the note. Does it identify the correct job? Does every flagged item exist? Does it preserve an unresolved answer? Can the reviewer find the evidence? Use the same criteria for manual and assisted work.

Choose test cases that represent the work you intend to automate. Include clear records, incomplete records, a changed order and an ambiguous request. Keep examples used to tune the system separate from examples used to judge it. Report the selection method so a good result on easy cases is not mistaken for coverage of the whole queue.

Record minutes where people actually spend them

Use one row per attempted task, including rejected outputs. A rejected draft still consumed preparation and review time, and the work may still need to be completed manually.

FieldWhat to record
Case and task typeA stable identifier and a description of the work
Accepted resultThe conditions the finished work must meet
Manual baselineHuman minutes to finish a comparable case
PreparationTime retrieving, cleaning or explaining source records
ReviewTime checking the proposed result
CorrectionTime repairing an otherwise usable result
FallbackTime completing the task manually after failure
OutcomeAccepted, corrected, rejected or unresolved
Failure reasonMissing data, wrong inference, wrong action or another specific cause

Record elapsed turnaround separately from human effort. A background task may take longer on the clock while consuming less staff time. Conversely, a fast response may keep a reviewer busy for several minutes. Both measures can matter, but they answer different questions.

Work through this illustrative comparison

Suppose a team tests 20 fictional readiness notes. Manual work takes 12 minutes per note, so the comparison baseline is 20 × 12 = 240 human minutes.

The assisted batch produces 15 notes accepted after review, three that need correction and two that require manual fallback. All 20 tasks eventually reach the same acceptance criteria.

Human work in the assisted batchMinutes
Prepare the records30
Review the proposed notes70
Correct the three repairable notes50
Complete two failed cases manually30
Total task effort180
Allocate ongoing maintenance for the period20
Total effort including maintenance200

For this fictional batch, the saving is 240 − 200 = 40 human minutes. Dividing by 20 completed tasks gives an average saving of two minutes per task. Before counting maintenance, the apparent saving would have been 60 minutes. The maintenance assumption changes the decision.

These numbers illustrate the calculation; they are not expected AI performance. They also do not include the initial setup effort. Track that separately and decide how it will be recovered over the useful life of the workflow. Do not hide it by excluding it from every future review.

Keep the failures visible when averages look good

An average can conceal a slow or consequential exception. Show the range of review times and inspect the cases that took the most effort. Separate a harmless style correction from an unsupported amount, wrong customer or unapproved action.

Record incorrect outputs that were caught and incorrect outputs that reached the next step. The two reveal different weaknesses. Also inspect good outputs that reviewers rejected; a system can create little value if the team cannot understand or trust its evidence.

A small pilot should not be described as a controlled causal study. Staff may work faster as they learn the cases, and the chosen cases may not match next month's queue. Use the results to make a bounded operating decision and keep measuring after that decision.

Decide the next step before you see the result

Write an acceptance rule with the person who owns the work. Specify the minimum useful saving, the errors that stop the pilot and the effort available to maintain it. Choose thresholds based on this task's consequences and economics, not a universal percentage copied from another company.

Then apply the rule:

  • Expand only the task types that meet the agreed conditions.
  • Narrow the workflow when one class of case drives the corrections.
  • Improve source records when the system repeatedly lacks information.
  • Stop when acceptable work takes more effort than the manual process or the unresolved risks exceed the agreed boundary.

Give each correction an owner and preserve the failed example. Retest it after the change, then try an unseen case. A fix that only handles yesterday's example has not established a general improvement.

Bring the ledger, the actual outputs and the slowest cases to the review. Those records support a decision about the next batch. If the intended result is still hard to judge, use the guide to choosing an agent task before expanding the pilot.

How we research and review these articles