A pilot becomes a selection failure when its apparent value depends on follow-on commitments that were undisclosed or misunderstood, defeating the buyer’s learning or exit objective. Commercial interest in follow-on work is not itself evidence of anything wrong. Nearly every supplier hopes a pilot leads somewhere.
The diagnostic
One question separates the cases: what useful evidence and permitted artifacts remain if the buyer purchases nothing further.
A pilot that leaves reusable evaluation sets, documented findings, exportable data, and a record of what was tested has produced value independent of the next transaction. One that leaves a hosted demonstration, an expiring concession, and no exportable artifact has produced a reason to buy.
Four specific things to look for: commitments that were not disclosed, exports that exist but are unusable, concessions that expire in a way that changes the economics of leaving, and terms that block independent evaluation of the result.
Each has a benign explanation that has to be tested first. Scoped discovery legitimately produces findings rather than assets. A managed service legitimately runs on the supplier’s infrastructure. A misunderstanding about deliverables is a scoping failure rather than a deception, and a low fee alone evidences neither. The buyer’s response to a genuine misunderstanding is to revisit scope or selection, not to infer intent.
Price is a question
Low prices have multiple explanations, and tender guidance requires giving a supplier the opportunity to demonstrate deliverability before a bid is rejected as abnormally low.
The same guidance warns that relative price scoring can work against value for money — a mechanism worth understanding, because it rewards the bid that is cheapest at the moment of scoring rather than the one that is cheapest to have.
Dependency is a trade, not a defect
Commercial dependence and technical dependence are different things, and portability trades against service benefits rather than dominating them. The recommended posture is monitoring switching costs and preparedness rather than eliminating every dependency, since elimination has a price that someone pays.
The practical instrument is a hosting or service business case that includes exit cost and time estimates for the period after discounts lapse, with those switching estimates reassessed by actually rebuilding or testing a component rather than by asserting them.
Paying for the comparison
A paid bake-off commissions bounded work from several suppliers to produce evidence for a decision. Funded competitions with phased down-selection are an established institutional form of this.
Payment buys comparability, and comparability is the deliverable. Agree the procurement route, compensation, deliverables, confidentiality, and rights before work starts. Give each supplier a defined task, common permitted inputs, stated resource limits, and disclosed constraints. Keep the specific assessment cases protected where appropriate while disclosing the evaluation dimensions and method.
Then evaluate artifacts and failures alongside demonstrations, and include a bounded exercise in which the receiving team operates, corrects, or hands over the result. Record access and assistance differences rather than silently crediting them to capability.
The recurring result is the one worth designing for: of two suppliers given identical inputs, the stronger presentation belongs to the deliverable that needs undocumented supplier intervention to run. Presentation quality and transfer evidence get recorded separately, under criteria disclosed in advance.
Evaluation discipline is the other half. Competitive proposal evaluation confined to disclosed factors, with documented strengths, deficiencies, weaknesses, and risks, is what prevents a comparison from being reconstructed after a preference forms.
Diligence on the transfer, not the build
Past performance is one indicator, and its relevance, currency, source, and context all bear on what it shows. Where relevant history is absent, neutral treatment is the correct response rather than an invented maturity score.
The evidence worth requesting is transfer-specific: a comparable transfer with its scope, recipient environment, and unresolved differences; a versioned artifact manifest with known omissions; actual licences, accounts, and third-party consents; and evidence of what a recipient demonstrably operated, corrected, or rebuilt, together with the assistance required.
Where that history is thin, a bounded permitted transfer exercise tests it directly. A vendor showing successful launches and a recipient exercise revealing a missing rebuild instruction have produced two facts, and only the second is about transfer.
The rule
What stays fixed is that a pilot’s value is measured by what survives the decision not to proceed. What changes is how much of that value is contractual and how much is practical, and both have to be checked, because permitted export and feasible export are different.
Not to be confused with
Vendor misconduct. The diagnostic identifies a structure, not an intent. No prevalence of the pattern is established, and a supplier that benefits from a follow-on sale has done nothing improper by benefiting.
A rule against managed services. Continuing supplier operation can be the right answer. It stops being a pilot at that point, and should be priced and compared as the service it is.