Accidental production is consequential use or reliance that has developed beyond the operating conditions the organisation explicitly accepted. The deployment’s name is not the test, and a formally authorised pilot using real data under suitable controls is not this failure.

What causes the transition is use. Somebody makes a real decision, about a real customer, that leaves a real record.

The single switch does not exist

Different requirements have different boundaries, so a change can cross one and leave another untouched. The workable instrument is a record of specific changes mapped to the authorities each one engages, rather than a prototype-to-production flag.

Seven events trigger reassessment: changed data access, real-user exposure, changed decision influence, changed intended purpose, changed geography, changed vendor dependence, and changed operational authority.

Each is answerable precisely. Who uses the system, which data enter it, what its output influences, and who can act on that output. Compare those against the last reviewed configuration and record the differences — without waiting for the project’s label to change, because the label is the last thing to move.

The characteristic sequence

An assistant is evaluated on synthetic questions by the team that built it. Later, staff connect member records and begin using its summaries while deciding actual applications. The project is still called a pilot.

Two things changed: data access and decision influence. Neither required an approval, because neither looked like a deployment. The name pilot proves no exemption, and equally proves no violation — the review identifies the institution and jurisdiction and checks each applicable requirement against the facts.

The variant worth watching for is subtler. Output presented as advisory becomes authoritative in practice, because the people receiving it stop disagreeing with it. Nothing in the architecture changed. Whether supposedly advisory outputs are relied on is a question about behaviour, and it belongs in the review.

What attaches, and when

Three regimes engage on different triggers, and it is worth seeing how differently they behave.

Third-party risk engages earliest. A business arrangement can exist without a contract and without payment, so the vendor relationship was already in scope before the data moved. Some control concerns require management and monitoring even after the relationship ends.

The EU framework distinguishes systems developed and put into service solely for scientific research from activity before placing on the market or putting into service, and expressly leaves testing in real-world conditions outside the second exclusion while preserving other Union law. Its definitions of putting into service include supply for the provider’s own use, and certain free supply counts as making available — so an internal tool given to staff at no charge is not obviously outside them.

Model risk management under current Federal Reserve guidance runs the other way for generative systems: footnote 3 of the SR 26-2 attachment excludes generative and agentic AI, directing the organisation to its own governance practices instead. Being outside that scope routes the work rather than ending it.

A supervisory position can also be that no AI-specific regulations have been issued while existing technology-neutral regulations apply and examiners review controls, compliance, monitoring, and vendor diligence. The absence of a named AI rule is not the absence of duties.

The retroactive part

This is what makes the failure expensive rather than merely embarrassing.

Once the system is inside a regime, the obligations attach to work already done. Validation evidence that was never produced, documentation that was never written, and adverse-action capability that was never designed for are all now missing from a record that describes real decisions about real people.

Records retention and discoverability change too. Interaction logs created during an exploratory phase, under retention rules chosen for an experiment, become the evidence trail for consequential decisions.

The prototype cannot go back and have been governed. It can only be stopped, remediated, or accepted with the gap recorded — and the third option requires someone with the authority to accept it.

The rule

What stays fixed is that scope follows the actual activity, the actual data, and the actual decisions. What changes is which of the seven events occurred, and each one is observable at the time by anyone who is looking.

The review’s precondition is an identified institution, a system and use description, actual data flows, and dated sources. Where those are unknown, the applicability question is recorded as unresolved rather than answered favourably by default.

Not to be confused with

Scope creep. Widening functionality during operate is a delivery problem. This is a governance problem, and a system can acquire regulated status without gaining a single feature.

A finding of violation. The boundary review determines which requirements apply and what evidence exists. Technical test success, budget approval, and legal authorisation are three separate determinations, and each has its own owner.