The operate phase exists to produce one thing: a recorded decision about what happens next, the evidence supporting it, and the person who owns the consequence. Extend, kill, and transfer name the outcomes. Three more hide inside them and deserve naming.
Six outcomes, not three
Transfer moves ownership to an accepting team. Kill ends the work. Extend continues it for a bounded purpose.
Pause stops activity while a dependency resolves, without concluding anything about value. Redesign keeps the problem and abandons the approach. Restricted continuation narrows the scope, the cohort, or the authority and continues within the smaller boundary.
Four distinct things also get conflated under these words and separating them prevents most of the confusion: ending the experiment, retiring the service, changing its scope, and transferring its ownership. A prototype can end while its service continues under someone else, and a service can be retired while the experiment that produced it is judged a success.
What each outcome requires
Transfer needs an accepting owner and demonstrated capability. Live-phase guidance asks for a sustainable operating arrangement, useful performance measurement, and support staff familiarity before a service moves into live operation, with improvement continuing afterwards. A positive evaluation satisfies none of those three. The evaluation says the thing works; transfer readiness asks whether anyone can run it.
Extension needs a bounded learning purpose. What question does the additional time answer, and what observation would end it? Extension without that is the default that happens when no decision was made, and it is the outcome most projects reach by not choosing.
Closure needs a managed ending and preserved lessons. The evaluation sets, the labelled data, the integration knowledge, and the record of what was tried survive the system. Losing them converts a completed experiment into a repeatable expense.
Retirement has its own grounds. Loss of user need, or a lack of cost effectiveness in the service’s current form, are both sufficient. Neither requires technical failure, and a well-functioning service whose need has gone should be retired rather than defended.
The default when evidence is missing
Inconclusive is a result, and it needs a response written before the review: close the test without a value conclusion, and decide explicitly whether another test is justified.
Without that route, inconclusive evidence flows into extension, because extension is the option requiring no one to defend a conclusion. The system keeps running, the budget keeps being spent, and no further evidence is gathered because nobody specified what would settle it.
Reconstructing the criteria is not available
The decision is made against the criteria as originally recorded, and the denominator is where that gets tested.
Take a plan requiring at least 90 confirmed successful outcomes from 100 eligible attempts. Telemetry reports 88 successes, 7 failures, and 5 unresolved. The declared condition is not met — there are 88 confirmed successes. Reporting 88 of 95 resolved cases as 92.6 percent has changed the denominator and answered a different question, and the five unresolved outcomes could contain successes or failures.
So the review inspects evidence against the original record, identifies unmet conditions, chooses an authorised response, and preserves the rationale. Where a criterion changed, both versions are retained with the reason, the timing, and who accepted it. Prespecified conditional rules are legitimate — an alternative analysis where an assumption fails, planned in advance — and planning alone does not prevent selective interpretation.
Population coverage stays visible alongside outcome rates. An undefined rate over a zero denominator is unavailable, not perfect.
What the operate phase should have been collecting
The decision is only as good as the instrumentation, and the instrumentation is designed backwards from the criteria rather than from what the system happens to emit.
Each criterion maps to events and fields: attempt identifiers, timestamps, release and metric versions, completion state, permitted evidence references. Successful, failed, interrupted, and duplicate paths are exercised in a test setting before the real run. Late arrivals, unavailable labels, retries, exclusions, and missing outcomes have declared treatments.
And what the system cannot observe is marked as such. An event count cannot include a request that never reached the collection point, so a clean telemetry record is consistent with a large invisible failure population.
The rule
What stays fixed is that the decision names an owner for whatever it produces — a receiving team, a closure, or a bounded next question. What changes is which of the six outcomes the evidence supports, and no outcome is the presumed destination.
Not to be confused with
A verdict on the idea. A stop decision and a conclusion about value are separate statements. The closure record keeps the evaluation result inside its tested scope and states why the work ended.
A single date. Extension, closure, and transfer each have their own readiness conditions, and reaching the end of a planned operating period satisfies none of them automatically.