A success criterion is falsifiable when a stated observation would end the project. It specifies an outcome claim and the observations that would support it, count against it, or leave it unresolved. A criterion that no achievable result could fail is not a criterion, however precisely its numbers are written.
The third state
Most success criteria have two states: hit or missed. Real evaluations produce a third result more frequently than either, and having nowhere to put it is the defect that does the most damage.
A usable criterion divides the outcome space three ways. There is evidence supporting a worthwhile benefit. There is evidence counting against a worthwhile benefit under the tested conditions. And there is a region where the observations cannot distinguish the two — because the sample was too small, the data were inadequate, the exposure was interrupted, or the uncertainty spans both a valuable improvement and no improvement at all.
Absence of observed benefit is not evidence of no benefit. Collapsing those produces confident failure reports about untested questions, exactly as a two-state criterion produces confident success reports about inconclusive ones.
What has to be written down first
Nine things, and the difficulty is that most of them look like details until the result arrives and the argument starts.
Claim and scope. The recipient, task, intervention, and conditions the conclusion will apply to.
Outcome. Its operational definition, unit, denominator, data source, and the treatment of missing or repeated observations. The denominator is where most disputes live.
Comparison. The baseline or alternative, and what licenses a descriptive rather than causal reading.
Practical importance. The smallest improvement worth acting on, with its rationale and the costs it must cover. This is the number that makes the rest meaningful, and it is set by what the change costs, not by what the instrument can detect.
Observation plan. Sampling, exposure, window, and the conditions that would make the data inadequate.
Analysis. Estimator, uncertainty, planned comparisons, and any multiplicity adjustment.
Monitoring. A fixed endpoint or a specified sequential procedure. Operational pause conditions are recorded separately, because stopping because something broke is not stopping because the answer arrived.
Interpretation. The three regions above.
Action and ownership. Who decides, and what each of the three results permits.
Statistical significance does not substitute for any of this. A p-value measures neither effect size nor practical importance, and a decision resting on a threshold alone has skipped the question of whether the effect is worth having.
Looking early
The rule that results must not be inspected before the end is wrong as stated. Sequential inference methods support continuous monitoring while controlling false positives, under the conditions those methods specify. What fails is unspecified peeking followed by stopping when the number is favourable, because the error control comes from the procedure being fixed in advance.
So the monitoring field is a genuine choice with two valid answers — a fixed endpoint, or a named sequential procedure with its own rules — and one invalid one, which is leaving it blank and improvising.
The rule
What stays fixed is that the criterion has to be capable of ending the work. What changes is the threshold, the outcome, and the window, all of which are project-specific and none of which can be copied from another project.
Where criteria fail
Two failures, from different causes.
The goalposts move after the result. A trial specifies at least 90 successful tasks out of 100, records 86, changes the threshold to 85, and reports the original target as met. Selecting among analyses and data-collection choices after seeing outcomes inflates false-positive rates, which is the finding the methodological literature established by simulation. The correction is not a ban on changes: preserve the original criteria, timestamp every revision, record which results were known at the time, and report the conclusion under both the original plan and the revised one. Fixing a verified instrumentation error is a legitimate change, and it survives that treatment intact. Lowering a threshold to clear it does not.
Discovery metrics harden into confirmatory ones. Refining a measure while still learning what to measure is proper work. The failure is relabelling a favourable exploratory analysis as the original test. Each refinement gets a new dated version, and once outcome data have been examined, that knowledge is disclosed as part of explaining the change.
Not to be confused with
A preregistration. A timestamped local file preserves a decision history and is worth keeping. It is not a public preregistration unless it was actually registered, and describing it as one overstates what anyone can check.
A large sample. Volume does not justify the comparison, the outcome definition, or a causal reading. It narrows one kind of uncertainty and leaves the design questions untouched.
An operational pause. Suspending an evaluation because access broke is an interruption, not a result. Reporting it as a disproved hypothesis converts a logistics problem into a false finding.