An agent built a Tauri desktop application end to end. The pure logic was correct on the first or second attempt. Every defect that mattered sat at a boundary no headless test can reach.
7 min read
A compressor reconstructed from behaviour alone, graded against an answer key sealed until the work was frozen. The decoder scored 100 percent and the encoder 8.3 percent on the same format.
7 min read
Verne's submarine crushes at 340 metres against a narrated 16,000. Extracting a falsifiable spec from fiction, and what it takes to trust numbers an AI produced when no external oracle exists.
7 min read
Characterization tests preserve selected observed behavior. Their baseline must distinguish required behavior from existing defects and comparison artifacts.
4 min read
Parallel replacement compares implementations over matched inputs and state before changing operational authority. Shadow execution must control external effects.
4 min read
Evaluating generated assertions through independent examples, approved baselines, and meaningful fault detection.
3 min read