How it works

Three gates decide whether a page can be used at all. Crawl access: separate user agents govern training, search indexing and live answer retrieval, and blocking one is not blocking the others. Extractability: passages that state a claim plainly, near a heading, with dates and figures in text rather than images, survive chunking better than material spread across a long narrative. Corroboration: claims repeated in independent, identifiable sources are safer for a system that must attribute what it says.

Because candidate retrieval is a search problem, ordinary ranking factors still apply. Because generation is sampled, the specific citations attached to an answer are less stable than a search result page.

Example

A pricing page that renders its numbers inside an image contributes nothing quotable, even when it ranks. A page with a short, dated, plainly worded pricing table is quotable — and quotability, not persuasion, is what gets a source into an answer.

Why it matters

Publishers and vendors are being asked to optimise for systems whose selection process is only partly documented. The defensible response is to control the parts that are documented — crawl access, structure, dates, authorship, primary evidence — and to measure the rest empirically, per prompt, over time, rather than trusting a single observation.