Statistical power is the probability that a specified test rejects its null hypothesis under a specified alternative and sampling model. Pilot cohorts are assembled from whoever is available, willing, and permitted. Those two facts rarely meet, and the gap between them is why so many pilots produce an inconclusive result that gets reported as a positive one.

The arithmetic that decides it before you start

Sample-size requirements depend on the error rates chosen, the shift worth detecting, and the outcome’s variability — with the calculation differing according to whether the standard deviation is known.

Work an example. Define improvement as baseline time minus later time, assume one normally distributed mean change per person for 12 independent people, a known population standard deviation of 10 minutes, a predeclared one-sided test at alpha 0.05, and target power of 0.80. The detectable shift is the sum of the two critical values multiplied by the standard deviation over the square root of N: about 7.2 minutes.

If the business case claims a 5-minute saving, that pilot cannot find it. Detecting 5 minutes under the same assumptions needs about 25 people.

This calculation is available before recruitment, and it is the one calculation that determines whether the exercise can answer its question. Running it afterwards, as post-hoc power, tells you nothing you did not already know from the result.

What breaks the calculation

The unit of assignment. Assignment by case, person, team, or period changes the analysis. Six weeks of repeated work by twelve people does not create hundreds of independent observations, because within-person dependence is real and unmodelled dependence inflates apparent precision.

Optimistic nuisance parameters. Planning with the most favourable variance estimate from a small pilot produces a study sized for a world that is not this one. Sensitivity to plausible values belongs in the plan.

Carryover. Where using a tool teaches a persistent skill, returning someone to the old interface does not restore their earlier state. An alternating schedule has to account for that, or the comparison measures learning.

Attrition, measurement error, and spillover. Each needs a stated assumption, because each moves the required N.

Baseline is not counterfactual

A baseline describes an observed starting situation. A counterfactual describes what would have happened under the alternative — continuing the existing workflow through the same later period. The baseline informs the counterfactual and does not substitute for it.

Take handling time falling from 40 to 30 minutes in the prototype group while a comparison group falls from 35 to 29. The prototype change is -10, the comparison change is -6, and the difference of differences is -4 minutes. Reading -4 as the effect requires a defensible design: the parallel-trend assumption, comparable measurement, and no unaccounted differential shock. Four averages establish none of those, and they establish nothing about uncertainty.

Anticipation matters too. Outcomes can move before formal rollout, so baseline collection is planned early rather than measured once the project is announced.

And a before-and-after comparison without a comparison group provides weak support for attributing change to the intervention. That is the design most pilots actually run.

Choosing a design that answers the question

Several alternatives exist and none is superior in the abstract.

Paired offline replay asks whether candidates meet specified properties on the same admissible cases, and leaves live operator behaviour unobserved. Shadow operation surfaces outputs and failures on real workload, and cannot demonstrate changed operator performance because operators never saw the output. Randomised phase or sequence designs address assigned timing and exposure, subject to clustering, spillover, and persistence across periods. Interrupted time series needs enough suitable periods and credible treatment of concurrent change, seasonality, and autocorrelation. Matched or trend comparison depends on selection and on the assumptions of the chosen method.

A shadow study generating summaries for 500 files while staff work unchanged supports a rubric assessment of the text and an inspection of operational failures. It has not observed anyone becoming faster. That claim needs a workflow study, and saying so is the difference between a bounded result and an overreach.

Where the effect meets the business case

The evidence chain has four links and they get collapsed: a measured change, an attributed effect, a valuation, and a realised impact.

Saved staff minutes are a measured change. They become cash only through a mechanism that reduces expenditure, and where salary spend is unchanged there is no such mechanism. Fewer contacts are not successful deflection until completion and repeat contacts are known. Faster processing is not better service until end-to-end elapsed time and the next team’s workload are assessed. Lower current spend can be delayed work or reduced output.

Reported efficiency savings account for delivery costs and avoid double counting, and where two initiatives contribute to the same result, contribution reporting and an additive total are different things. An equal split is an assumption unless justified.

The disposition for each contested claim is one of four: accept within scope, narrow or reclassify, retain as a forecast, or withhold pending evidence. A supported non-cash operational benefit is retained as what it is rather than discarded for failing a cash test it was never making.

The rule

What stays fixed is that the minimum detectable effect is a property of the design, computed before recruitment. What changes is the claimed benefit, and where the claim is smaller than the detectable effect, the pilot’s aim is narrowed to feasibility and the evidence that would settle value is named.

Not to be confused with

A prohibition on small studies. Small pilots are useful for implementation problems, uptake, and variability. The error is using one to answer a question about magnitude.

Certainty from randomisation. Assignment does not produce perfectly identical groups. Realised imbalance and design failures still need attention.