Agree before testing

QuestionOur answer
What specific work are we testing, and what is outside the scope?…
Who leads the pilot, and who decides whether to continue?…
Who accepts the output, and what requirements must it meet?…
What data and tool actions are allowed?…
What must improve for the next step to make sense?…
Which errors can be corrected, and which require a stop?…
What routine and difficult cases will we include?…
Which cases will we reserve for the final evaluation?…
What are the scope, budget, and deadline for this test?…

Record a single case

Use minutes for time entries. For the AI-assisted method, include every attempt and any manual completion. Record roles and time without counting the same work twice.

ItemCurrent processWith AI
Case identifier, type, and difficulty……
Input preparation, tool operation, and drafting……
Review and corrections, including failed attempts……
Handoff and later work after the recipient returns the output……
Total human work……
Breakdown of total work by role……
Elapsed time from request to acceptance, measured start to finish……
Result: no correction / after correction / not accepted……
Reason for return, error severity, and resolution……
Model, instruction, and source versions; test date—…
Tuning / final evaluation; repeat of the same case……

If someone watches generation, include that time in tool operation. An unattended model run is waiting time. When work overlaps, do not add the individual durations to calculate delivery time.

Summary for leadership

Decision or evidence neededOur answer and a link to the evidence
Recommendation: expand / revise and test / stop…
Number of distinct cases and number of repeat runs…
Coverage of routine work, exceptions, and gaps…
Quality of both methods: no corrections / corrected / not accepted…
Average, median, and most demanding cases by task type…
Human work for both methods over a comparable period…
Difference in time to deliver the result…
Work volume and assumptions for extrapolating to a month…
Recurring maintenance, responsible role, and time over the same period…
Tool price and other recurring costs over that period…
One-time preparation to date, separate from further investment…
Cost of the proposed next step, including people’s time…
Who saves time, who takes on work, and how will we use the capacity?…
Whether the agreed conditions are met and which risks remain…
What do we not know, and how could it change the recommendation?…
Next step: scope, owner, budget, date, and stopping condition…

Check the calculation: subtract all human work in the AI-assisted method and its additional recurring maintenance from human work in the current process, over the same period. If the current method has its own maintenance, include that too. A monthly projection only holds for the corresponding volume and mix of work. Minutes from different roles do not represent identical monetary costs. Do not label freed-up time as payroll savings without further evidence.

A short recommendation to send to leadership

Decision: …

Verified: … across … distinct cases; repeat runs …

Benefits, costs, and risks: …

Still unknown: …

Recommendation: … because …

Next step: …; owner …; budget …; decision date …; stop if …