Text written for retrieval is read in fragments. A chunk arrives without the paragraph that preceded it, without the section heading that framed it, and without the sentence that introduced the term it uses.

That is a constraint on how arguments are built, not on how they are formatted. Prose style is discarded by the process; the structure of the claims is what survives or fails to.

What a fragment has to carry

A claim that depends on its surroundings loses its meaning when separated from them. So the unit of composition becomes a sentence or short passage that names its own subject, states its own relation, and carries its own qualification.

Three things go first under chunking, and they are the three that make prose pleasant: transitions, callbacks to earlier material, and context accumulated across paragraphs. A sentence beginning “this means that” has no referent in a retrieved fragment.

Three things survive: a definition stated once in full, a distinction between two named things, and a claim with its own scope attached.

The practical consequence is that qualification cannot live at a distance. A claim qualified three paragraphs earlier arrives unqualified, which is how a bounded finding becomes a general one — the qualification stayed behind.

Tables are the extreme case

A table loses its header row at a chunk boundary, and the cells become uninterpretable strings.

Worse, a table asserts parallelism. Rows in identical columns claim that the things in them are the same kind of object, which is the claim the entry should be examining rather than assuming. Writing the rows as sentences forces each one to state its own subject and relation, and the ones that resist are the ones that did not belong in the table.

The mirror problem confirms it from the other direction: a conversion that drops a column labelled illustrative produces output that parses correctly and has lost the qualification that made the content honest.

One concept, one address

A canonical page for each distinct reference question is the structural counterpart. Redirects, canonical annotations, and sitemap inclusion are signals for choosing among duplicate or similar pages, and where the concept boundary falls is an editorial judgement rather than a rule.

Two failures bracket it. Splitting one concept across pages means no fragment contains the whole definition. Pointing substantively different topics at one URL, to tidy a sitemap, means retrieval cannot distinguish them.

Stable topic identifiers map to public addresses independently of the current category path, so renaming a term or reorganising a hierarchy does not break the record. An alias register preserves moves, and the address of the current entry stays distinct from an immutable reviewed version — changing a name should not rewrite historical attribution.

The reader check is direct: arrive from a link with no preceding context, identify the referent, the scope, the evidence date, and the unresolved limits, follow a source, and return.

Structured data validates syntax, not meaning

Linked-data representations map compact terms to identifiers, and a defined-term vocabulary supplies name, description, term code, and containing set.

Three checks are separate and only the third is about correctness: JSON syntax, vocabulary expansion, and semantic comparison against the visible content.

An exporter converting a proposed readiness check into a readiness certification produces valid JSON. The syntax passes and the claim has been broadened, which is the failure mode of every automated representation — it preserves shape and drops qualification, exactly as chunking does.

So provenance, claim status, qualifications, and version history need a reviewed mapping rather than a default one, project-specific fields do not go silently into a public vocabulary namespace, and navigation links do not become assertions of semantic equivalence.

What none of this establishes

An agent-oriented site index is an optional interface to test against intended consumers. File presence, directory listings, and crawler access establish neither accurate retrieval nor citation nor training inclusion, and at least one search provider states that its AI features require no new AI text files or special markup, while noting that meeting eligibility requirements does not guarantee indexing or serving.

A server log recording a fetch establishes a request. It does not establish successful discovery, and it certainly does not establish that the answer cited the right term.

Measuring adoption therefore requires holding content and version constant and inspecting retrieval and answers separately. Availability and effect are different observations, and the format work is worth doing on the argument that fragments should be self-contained rather than on a promised ranking outcome.

The rule

What stays fixed is that every claim carries its own scope. What changes is how much surrounding context a given reader has, and writing for the reader with none is what makes the text work for the reader with all of it.

Not to be confused with

Search optimisation. No ranking or citation outcome follows from any of this. The argument is about which claims survive fragmentation, and it would hold if no search engine existed.

Simplification. Self-contained claims are frequently longer than context-dependent ones, because the qualification has to travel with them. Compression at the paragraph level and survival at the fragment level pull in opposite directions.