
Imagine leadership reviewing whether to give more people access to AI. The team brings two commentaries on the same sales report. A person spent 30 minutes of active work on the first; the model generated a first draft of the second in two minutes. A slide that shows “30 minutes versus two” looks conclusive. But those figures measure two different things.
An AI-generated draft is not yet something the business can act on. Someone still has to check the numbers, trace the sources, correct the conclusion, and answer questions from the person receiving the commentary. Count that work, too.
Before recommending further investment, I would want to see the whole process, from the initial request to an accepted, usable commentary. I would also want to know what the business will do with any capacity it frees up. A successful pilot can give you a good reason to continue. Learning that further investment does not make sense under current conditions can be just as valuable.
What you need to know before investing more
I would start with three questions. Will the output meet our quality requirements? Will it free up capacity that people can actually use? Can we operate the new process at an acceptable cost and level of risk? If an answer is missing, the team should be able to explain what still needs to be tested.
For this example, a commentary on a sales report must do more than read clearly. It needs the right reporting period and revenue definition, accurate figures, and a clear distinction between a documented reason for a change and a hypothesis. The person receiving the commentary needs to know whether they can rely on it when making a decision.
Write down the quality criteria before testing, then apply them to both methods. If people already disagree about what revenue includes, first clarify the meaning and reliability of the company’s numbers. Switching models will not resolve that disagreement.
Where the promised time savings went
The following numbers are a model example, not results from one of my projects or a recommended benchmark. Assume both processes produce an accepted commentary that meets the same quality requirements. The table counts active human work only. In the current process, the first row covers preparation and drafting; with AI, it covers the human time spent preparing the prompt. The model’s two-minute generation run is measured separately.
| Active human work per commentary | Current process | With AI |
|---|---|---|
| Preparing a draft / preparing the AI prompt | 20 minutes | 5 minutes |
| Review and corrections | 8 minutes | 15 minutes |
| Handoff of the result | 2 minutes | 2 minutes |
| Total | 30 minutes | 22 minutes |
In this example, AI speeds up the first draft, but review takes longer. The difference in human work is eight minutes per commentary. The two minutes of unattended generation are not in the table: the person can do something else while the model runs. If someone has to watch the process, include that time in their work.
At 100 comparable commentaries per month, the gross difference would be 13 hours and 20 minutes. Add another model assumption: the new process requires two hours of additional recurring maintenance per month. That leaves 11 hours and 20 minutes of capacity for other work. The estimate still excludes the tool’s price, one-time setup, and any further investment needed to expand its use.
To understand when the recipient gets the result, also measure elapsed time from the initial request to acceptance. That includes waiting. When activities overlap, you cannot calculate it simply by adding the entries in the table.
Saved time needs a useful destination
The total hours saved will not tell you whether to expand the pilot. It matters whose time is freed up, which tasks it comes from, and what that person can do instead.
If a salesperson can use the time to prepare proposals that are currently being postponed, the business has a reason to keep testing that benefit. If the saving is spread across many people in a few minutes here and there, with little practical change in their work, the same value will be harder to demonstrate. A time sheet alone cannot tell you which situation you have.
Look at how the work moves between people, too. If a junior employee prepares the draft and your busiest analyst takes on the corrections, the total minutes can fall while your capacity bottleneck gets worse. Ask the team to show who saves time and who takes on more work.
Freed-up capacity does not automatically reduce payroll spending. A financial assessment needs the operating cost, the time required from each role, and a defensible account of the value of their other work. Separate one-time setup from recurring costs, and compare costs and benefits over the same period. Record the work already invested, but justify the next budget by what it can deliver from this point forward.
What the team should provide for a credible comparison
Leadership does not need to take over technical testing. It does need to understand what the results cover. I would ask the person leading the pilot for five things:
- A defined scope and acceptance criteria. What type of work is being tested, who accepts the result, and how do they identify an error? Which errors can be corrected, and which require an immediate stop? Agree on the data and actions the tool is allowed to use.
- A comparison covering routine and difficult cases. Show how the cases reflect actual work, including incomplete inputs and common exceptions. A person repeating the same task already knows the answer; alternating the order of the methods or using tasks of comparable difficulty can help.
- All human work, including failed attempts. Count preparation, tool operation, review, corrections, handoff, and errors discovered by the recipient. If someone completes the task manually after three failed AI attempts, include the entire sequence. Record the roles involved and track delivery time separately.
- An evaluation on additional cases after tuning ends. Results on tasks used to refine the instructions may be overly favorable. Reserve some cases for the final evaluation. Record the model, instruction, and source versions, and distinguish any changes made during the comparison.
- The range of outcomes and the limits of the conclusion. Alongside the average, show the number of cases, acceptance without corrections, median time, and the most demanding corrections. Separate routine tasks from exceptions. Choose the number of cases based on the variety of the work and the consequences of errors; a small test does not establish reliability across the entire operation.
If the same input produces different outputs, repeat some of the difficult cases. Do not count those attempts as new, independent business situations. A decision requires a clear account of what was tested, not just a larger number on a slide.
Expand, revise, or stop
Write down the conditions for continuing in advance. The required savings, acceptable amount of rework, and risk limits will differ between an internal working draft and a document sent to a customer. A universal percentage would conceal those differences.
Expand when the pilot meets the agreed conditions, recipients can use the output, and the benefit justifies the additional cost. Approve a defined next scope, an owner, the necessary resources, and a date for another review. Results from one process do not justify deploying it everywhere.
Revise and test again when you understand the cause of a problem and can define a further test. In our example, a missing link to the source of a number might be making review slower. You could test whether adding that source reduces review time while maintaining quality. The next attempt needs a clear question, a limited budget, and an agreed stopping condition.
Stop when corrections consume the benefit, nobody uses the result, or an unacceptable error occurs. The pilot has supplied evidence for a decision before a larger investment. Expansion does not have to be its outcome.
What a recommendation to leadership could look like
Our illustrative calculation does not yet support a broad rollout. We have an estimate of available capacity, but not all the costs or results from a representative set of tasks. An honest decision brief might read:
Decision: Whether to expand the use of AI for preparing sales report commentaries.
What the example suggests: At the same quality level, human work would fall from 30 to 22 minutes. At 100 commentaries and two additional hours of monthly maintenance, that would free up 11 hours and 20 minutes.
What is missing: Validation on routine and difficult tasks, the distribution of work by role, the cost of the next step, and an agreement on how the saved time would be used.
Proposed next step: Do not expand yet. Define a test to obtain the missing evidence, with an owner, scope, budget, and decision date agreed in advance.
You can use the team worksheet for your own pilot. It includes a decision-brief template and a record for each test case. Filling it in also makes the remaining unknowns visible.
This approach fits bounded tasks whose output can be checked. Time measurements alone are insufficient for rare, serious errors or effects that only become visible months later. And if AI enables work the business has never done before, assess the value of that new output; comparing it with a nonexistent previous process would not be meaningful.
If leadership and the team disagree about what the pilot has demonstrated, it can help to review the process, criteria, and evidence together. That is the basis of my AI workflow sprint. Tell me what work you are testing and which decision you are facing. That gives us a starting point for deciding what to address next.
Methodological background: Anthropic discusses evaluating AI agents, including success criteria, variation between runs, and checking actual outcomes. The capacity calculation and leadership brief in this article are a proposed approach, not an adopted standard or a measured case study.