# AI pilot: a decision brief

A worksheet for “How to evaluate an AI pilot before investing more.” Copy it into your own document. Where evidence is missing, write “not verified” and name the person who will obtain it. Completing the form does not replace testing.

## Agree before testing

| Question | Our answer |
| --- | --- |
| What specific work are we testing, and what is outside the scope? | … |
| Who leads the pilot, and who decides whether to continue? | … |
| Who accepts the output, and what requirements must it meet? | … |
| What data and tool actions are allowed? | … |
| What must improve for the next step to make sense? | … |
| Which errors can be corrected, and which require a stop? | … |
| What routine and difficult cases will we include? | … |
| Which cases will we reserve for the final evaluation? | … |
| What are the scope, budget, and deadline for this test? | … |

## Record a single case

Use minutes for time entries. For the AI-assisted method, include every attempt and any manual completion. Record roles and time without counting the same work twice.

| Item | Current process | With AI |
| --- | --- | --- |
| Case identifier, type, and difficulty | … | … |
| Input preparation, tool operation, and drafting | … | … |
| Review and corrections, including failed attempts | … | … |
| Handoff and later work after the recipient returns the output | … | … |
| Total human work | … | … |
| Breakdown of total work by role | … | … |
| Elapsed time from request to acceptance, measured start to finish | … | … |
| Result: no correction / after correction / not accepted | … | … |
| Reason for return, error severity, and resolution | … | … |
| Model, instruction, and source versions; test date | — | … |
| Tuning / final evaluation; repeat of the same case | … | … |

If someone watches generation, include that time in tool operation. An unattended model run is waiting time. When work overlaps, do not add the individual durations to calculate delivery time.

## Summary for leadership

| Decision or evidence needed | Our answer and a link to the evidence |
| --- | --- |
| Recommendation: expand / revise and test / stop | … |
| Number of distinct cases and number of repeat runs | … |
| Coverage of routine work, exceptions, and gaps | … |
| Quality of both methods: no corrections / corrected / not accepted | … |
| Average, median, and most demanding cases by task type | … |
| Human work for both methods over a comparable period | … |
| Difference in time to deliver the result | … |
| Work volume and assumptions for extrapolating to a month | … |
| Recurring maintenance, responsible role, and time over the same period | … |
| Tool price and other recurring costs over that period | … |
| One-time preparation to date, separate from further investment | … |
| Cost of the proposed next step, including people’s time | … |
| Who saves time, who takes on work, and how will we use the capacity? | … |
| Whether the agreed conditions are met and which risks remain | … |
| What do we not know, and how could it change the recommendation? | … |
| Next step: scope, owner, budget, date, and stopping condition | … |

**Check the calculation:** subtract all human work in the AI-assisted method and its additional recurring maintenance from human work in the current process, over the same period. If the current method has its own maintenance, include that too. A monthly projection only holds for the corresponding volume and mix of work. Minutes from different roles do not represent identical monetary costs. Do not label freed-up time as payroll savings without further evidence.

## A short recommendation to send to leadership

> **Decision:** …
>
> **Verified:** … across … distinct cases; repeat runs …
>
> **Benefits, costs, and risks:** …
>
> **Still unknown:** …
>
> **Recommendation:** … because …
>
> **Next step:** …; owner …; budget …; decision date …; stop if …
