The most useful generative AI task may not be the most impressive demonstration. In its evaluation of a Microsoft 365 Copilot trial, Treasury reported that people found the tool helpful for some basic administrative work, while more complex tasks exposed limitations.
From quick trial to useful evidence
The finding behind the headline
Treasury's February 2025 evaluation says almost all trial participants used Copilot at some point, but only 22 percent reported using it four to five times per week. Almost two thirds found it appropriate and beneficial for basic administrative tasks and work processes. The report also notes that value for complex tasks was limited by product functionality and privacy or security restrictions.
That distinction matters. A trial can show that a tool is available and interesting without showing that it is suitable for every role, dataset or decision.
Why small tasks make good pilots
A routine first draft, a meeting summary or a document search has a visible beginning and end. You can compare effort, quality and rework with the existing method. A high stakes task with vague success criteria gives you a weaker answer and a larger risk surface.
Start where the existing information is already approved for that environment. Measure whether people accept, edit or discard the output, not just how quickly a button produced it.
A test worth running
Take one recurring internal update. Have the team prepare it as usual for two weeks, then try an AI assisted first draft for two weeks. A reviewer scores factual accuracy, completeness and editing time without knowing which draft came from which method.
- Use the same source material for both methods.
- Track corrections and review time, not just time to first draft.
- Record cases where the tool made work slower or less clear.
Want to explore how this could apply to your organisation?
Explore Private AI ↗