Three AI workflow ideas, one team: Decide which pilot earns the time
Compare competing AI workflow ideas using evidence, review capacity and reversible actions. Includes an elimination sheet and an illustrative workload calculation.
When several AI ideas compete for the same team, compare the evidence and the review work each would require. Exclude candidates with unresolved data access, unbounded actions or no person able to judge the output. Then choose the eligible candidate that can answer the most useful question within the capacity you actually have.
This is a portfolio decision. If you have not yet defined an individual task's input, output and approval boundary, start with choosing a first agent task. The worksheet here compares several already defined candidates; it does not repeat that task-design exercise.
Put three proposals on the same page
An invented operations team is considering three pilots:
- Draft a summary of a completed service visit for the office to review.
- Prepare proposed invoice lines from approved work records.
- Write a reply to a customer's technical question using approved product documents.
Each sounds useful. The useful comparison is what a reviewer must inspect, what could go wrong, and whether the team can learn from a small test. Do not assume the most frequent task deserves the first pilot. A high-volume task with an expensive review step may consume the capacity you hoped to free.
NIST's voluntary AI Risk Management Framework emphasizes understanding context, measuring risks and managing them throughout use. The comparison below is an original DATUM planning method consistent with that approach, not a NIST assessment or certification. NIST AI RMF Core
Apply exclusions before assigning preferences
Use a yes, no or unresolved answer for these questions. Do not convert an unresolved answer into a low score that a large predicted benefit can outweigh.
| Entry condition | Evidence needed | If missing |
|---|---|---|
| Permitted inputs are identified | Exact fields, records and access owner | Resolve access before testing |
| Output has a judge | Named role and a reference for a correct answer | Define acceptance or exclude |
| Actions stay within a boundary | Draft destination and blocked actions | Reduce scope |
| Errors can be caught before consequence | Review step and escalation route | Redesign the workflow |
| Review capacity exists | Available time and backup owner | Reduce sample or defer |
| A baseline can be observed | Current task time and error evidence | Measure before estimating gains |
Some tasks may remain eligible after scope changes. A proposal to send invoices might become a draft-only review pilot. A technical-answer pilot may be narrowed to a single approved manual and questions whose answers appear in it. Record that narrower scope in the candidate name so a later reviewer does not approve the broader idea by mistake.
Compare eligible candidates on their real demands
Fill the table with your own observations. The entries below are illustrative assumptions, not findings about these workflows in general.
| Candidate | Available evidence | Main review burden | Learning the pilot should produce |
|---|---|---|---|
| Visit-summary draft | Approved test notes and a required summary format | Check omissions against source notes | Can the draft preserve the facts the office needs? |
| Invoice-line draft | Authorized scope, rates and completion records | Reconcile quantities, prices and missing evidence | Can draft preparation save time after reconciliation? |
| Technical-reply draft | Approved manuals with version IDs | Check the answer and confirm source applicability | Can unsupported questions reliably reach a person? |
If two candidates are otherwise suitable, prefer the one with accessible evidence and an available reviewer. That is a practical tie-breaker, not a prediction that it will deliver the largest long-term return.
Write down what makes the other candidates wait. “No approved document owner” is actionable. “Lower AI readiness” is not.
Account for the review bottleneck
Suppose the team's planning allowance is 90 reviewer minutes per week. An invented invoice pilot would review 15 drafts at four minutes each and reserve 20 minutes for investigating exceptions. Its planned demand is 15 × 4 + 20 = 80 minutes. That leaves 10 minutes of headroom.
A second pilot needing 40 reviewer minutes cannot run alongside it within that allowance: 80 + 40 = 120, which exceeds capacity by 30 minutes. The team must choose, reduce samples, or explicitly allocate more capacity. Calling both pilots “small” does not change the arithmetic.
These are planning assumptions. Time the first reviews and update the allowance. Count the work of preparing inputs and resolving disagreements as well as checking outputs. The AI pilot measurement worksheet shows how to compare the total effort with the current process.
Leave with one decision and two documented deferrals
The decision record should name the selected candidate, its permitted actions, reviewer, sample, baseline, stop condition and review date. For each deferred candidate, name the missing evidence or capacity that would make it eligible again.
A pilot can earn its place by exposing a constraint, including a result that says the task should remain manual. What matters is that the test produces a usable decision without silently taking time from the work it is meant to improve.