A scanned NASA specification, a novel, the memory of a product, and a sealed algorithm. What four controlled rebuilds establish about where AI-assisted reconstruction of legacy systems actually fails.
8 min read
A compressor reconstructed from behaviour alone, graded against an answer key sealed until the work was frozen. The decoder scored 100 percent and the encoder 8.3 percent on the same format.
7 min read
HAL/S runs again: twelve preserved Shuttle-era programs, verified value-for-value against an independent interpreter. The load-bearing human work was refusing a majority vote between OCR engines.
6 min read
Verne's submarine crushes at 340 metres against a narrated 16,000. Extracting a falsifiable spec from fiction, and what it takes to trust numbers an AI produced when no external oracle exists.
7 min read
A fixed seed makes a best effort, not a guarantee. Acceptance shifts from exact output to declared task properties, a predeclared sampling rule, and uncertainty measured at the unit actually sampled.
4 min read
Month-one and month-nine measurements are different quantities. A gate reading the early number transfers a system whose case has already expired, and time released is capacity rather than cash.
5 min read
Forecast 24 percent faster, felt 20 percent faster, measured 19 percent slower. The finding is about self-reported productivity as an evidence class, and it disqualifies most of what pilots collect.
5 min read
Killing a prototype destroys the system and preserves the encoded definition of correct behaviour. Salvage is a forward comparison, not a rescue justified by what was already spent.
4 min read
A success criterion is falsifiable when a stated observation would end the project. Most criteria have two states and need three, because inconclusive is the result that actually occurs.
5 min read
Three purposes, not three folders. One case can carry all three tags and still be one case family, and three passing tags are not three independent successes.
5 min read
A judge is a measurement component with order sensitivity and no established authority. Swap the presentation order, and a judge that picks the first answer both times has produced two conflicting verdicts.
4 min read
Every de-identification transform removes signal the model was going to use. What a pilot on treated data can and cannot establish, stated as a rule rather than a caveat.
5 min read
When three attempts cost what one used to, the gain is in comparison. The scarce resource becomes the ability to choose between working options, which is an evaluation problem.
5 min read
Cohorts are sized for access, not for power. Where the minimum detectable effect exceeds the claimed benefit, the pilot cannot produce the evidence its business case assumes.
5 min read
Three intervention points, not three exclusive strategies — the original RAG paper already combined two of them. Diagnose the failure before choosing the technique.
5 min read
Running a candidate against live inputs with its outputs withheld. The word does not establish the boundary — shared caches, writable destinations and quotas escape it unless someone checked.
5 min read
Not four rungs of one ladder. Three are construction choices that overlap and one is an operating context, and synthetic does not mean anonymous.
5 min read
The system can be rebuilt; the encoded definition of working cannot. Ship the harness first, and test it by making a deliberately wrong answer fail.
4 min read