A large trial is useful because it reveals patterns at scale. It is also easy to flatten its results into a single productivity claim. Australia's public service evaluation is more nuanced: adoption was moderate, use clustered around summarising and rewriting, and the gains were perceived rather than a universal time saving guarantee.
What the trial asked teams to separate
The reported numbers
The evaluation says one in three trial participants used Copilot daily. Summarising and rewriting were the most common functions. Sixty nine percent of survey respondents agreed it improved task speed, while 61 percent agreed it improved work quality. Around 40 percent said they could redirect time toward higher value work.
These are survey findings, not instrumented proof that every participant saved a particular number of minutes. The report also found up to seven percent of participants said Copilot added time to activities.
The hidden work is verification
The evaluation found that inaccuracy and lack of contextual knowledge meant people had to verify and edit outputs. That effort sometimes reduced the time saved. Access problems and weak functionality in some applications also shaped usage.
This is a practical lesson for any business: include review effort in the measure. A fast answer that needs extensive checking may not be a fast workflow.
Build a local scorecard
For a small pilot, measure adoption, task completion, editing time, material errors and user confidence separately. Sample the same task before and after. Keep a stop rule for sensitive or high consequence work and revisit the results when the model or product changes.
- Do not turn perceived time savings into promised ROI.
- Check who benefits and who receives extra review work.
- Reassess after major product changes.
Want to explore how this could apply to your organisation?
Explore Private AI ↗