An agent built a Tauri desktop application end to end. The pure logic was correct on the first or second attempt. Every defect that mattered sat at a boundary no headless test can reach.
7 min read
A scanned NASA specification, a novel, the memory of a product, and a sealed algorithm. What four controlled rebuilds establish about where AI-assisted reconstruction of legacy systems actually fails.
8 min read
A compressor reconstructed from behaviour alone, graded against an answer key sealed until the work was frozen. The decoder scored 100 percent and the encoder 8.3 percent on the same format.
7 min read
HAL/S runs again: twelve preserved Shuttle-era programs, verified value-for-value against an independent interpreter. The load-bearing human work was refusing a majority vote between OCR engines.
6 min read
Verne's submarine crushes at 340 metres against a narrated 16,000. Extracting a falsifiable spec from fiction, and what it takes to trust numbers an AI produced when no external oracle exists.
7 min read
When generation cost falls toward zero, the constraint relocates to verification, approval, and integration — none of which got cheaper. A queue forms where nobody is measuring.
5 min read
Understanding is measured on the receiving side, against named tasks. A team that passes acceptance and cannot explain why a rule change moved an outcome has not received the system.
5 min read
Cheap code lowers the cost of building the replacement. It does not lower the cost of being wrong about the legacy behaviour, and sent messages do not un-send.
4 min read
Forecast 24 percent faster, felt 20 percent faster, measured 19 percent slower. The finding is about self-reported productivity as an evidence class, and it disqualifies most of what pilots collect.
5 min read
When three attempts cost what one used to, the gain is in comparison. The scarce resource becomes the ability to choose between working options, which is an evaluation problem.
5 min read