A measurement taken in month one and a measurement taken in month nine are different quantities. Transfer gates read the first, business cases assume the second, and the difference between them decides whether a system arrives with a live justification or an expired one.

Four mechanisms, one curve

Early enthusiasm followed by decline is compatible with several explanations, and the trajectory alone diagnoses none of them.

Novelty response concerns reaction to unfamiliar technology and changes with experience. Observation effects concern behaviour altered by being studied — and the guidance is explicit that repeated interviews can themselves affect outcomes, trading against poorer recall if collection is delayed. Learning improves performance over the same period, pushing the other way. And support intensity is at its peak at the start.

Take a team measured as faster in week one and slower by week four, where week one also had daily specialist help and week four introduced harder files and a new model version. Four changes, one number, no isolation of any mechanism.

So the diagnostic is to record when each participant first received the intervention, their actual exposure, the support and observation contacts, the product version, and the workload. Then examine time since exposure separately from calendar time, because a later cohort meets different seasonal conditions and different staffing.

Withdrawals and non-use are preserved. Analysing only the enthusiastic survivors produces a stable favourable average from a shrinking population.

Time released is capacity, not cash

Three quantities get merged, and separating them is most of the work.

Task-time saving is reduced staff effort for a declared unit of work. Usable capacity is effort that can actually be reassigned under the operating constraints. Cash saving is an evidenced reduction in relevant expenditure.

The accounting is unforgiving. For a comparable workload, net minutes avoided are gross minutes avoided less the additional staff minutes — and those additional minutes must be genuinely additional to the baseline and absent from the gross comparison, or they get deducted twice.

Take 1,000 comparable episodes with 2,000 gross handling minutes avoided, where additional review consumes 400 minutes and additional repeat-contact work consumes 200. Net effort avoided is 1,400 minutes, or 23.3 hours. At 60 per hour that is 1,400 of attributed capacity value. With payroll unchanged it is not 1,400 of cash, and a new 300 tool bill is cash cost moving the other way.

A negative net result stays visible rather than being floored at zero. And where volume or task mix changed, effort is compared on a common basis before any time difference is multiplied by a volume.

Reported gains from published research need the same care. A study reporting a 15.2 percent increase in resolved chats per hour in its preferred model has reported that, for that population and that model specification. It has not supplied a conversion rate into hours or currency for a different organisation.

Deflection is three measurements and one claim

Consider 1,000 eligible episodes with complete follow-up, where 600 start without human involvement. Of those 600, 450 complete verifiably with no later human contact, 50 later reach a human, and 100 have no observed contact and an unknown outcome.

Initial no-human involvement is 60 percent. The full-window no-human-contact share is 55 percent. Verified self-service completion is 45 percent. Three different numbers, all correct, all routinely reported as the same thing.

None of them is 450 deflected contacts, because the baseline propensity to contact a human is unknown. Deflection is a causal claim; these are observations. Non-contact can also mean abandonment or contact through an unobserved channel, and unlinked users are not resolved users.

The one-year review

The review asks what the accepted system contributed during a declared period after transfer, with the starting event and cutoff stated and an explicit judgement about which benefits could mature by then.

It retrieves the original problem, baseline, benefit owner, assumptions, and forecast, and preserves subsequent scope changes rather than overwriting the case. It assembles actual use, task outcomes, total effort, operating costs, maintenance, incidents, displacement to other teams, and user burdens — with stated coverage, including non-users and abandoned workflows.

Take 1,000 tasks with an estimated three-minute reduction each: 3,000 minutes, 50 hours of released handling capacity. With staffing and paid hours unchanged, those 50 hours are not a payroll saving. What the capacity was used for, and what additional review and support work appeared, are the questions that decide the value.

Uncertainty in volume, unit effects, and attribution is preserved rather than multiplied into a precise return.

The rule

What stays fixed is that a benefit claim carries the duration and operating conditions it was observed under. What changes is both — the support tapers, the cohort widens, the case mix hardens — so the evidence required is evidence relevant to the period after transfer, not the period during the pilot.

Not to be confused with

A fixed observation window. No duration is established as sufficient. Four weeks, six weeks, and twelve months are all arbitrary without a stated mechanism and decision horizon.

Durability. A plateau during a short favourable period does not establish that the effect survives a different team, a different workload, or the withdrawal of pilot support.