Articles — Can AI read your site

Can AI read your site

AI often reads only part of your page

An index stores pieces of your page rather than the whole page, so your page can be linked while the price, the offer, or the useful part never reaches the answer.

Illustrative example, not a logged test. Northloop is a fictional example brand on this site.

A product URL for Northloop Trail Watch shows up as a source in the answer.

The paragraph next to it is about 30-day returns and warehouse dispatch.

What the person actually asked was who the product is for, and where to buy it in their country.

The system pulled one slice of your page, the shipping boilerplate, and wrote from that. Who the watch is for, the price, and the shop links never made it into the prompt.

This is a different miss from cited versus named. A citation is your URL used as a source. Being named is the product appearing as a pick. That split is a different article. Here your URL really is in the answer. The problem is that the wrong section of the page got used.

What a chunk is

Useful slice

“Who it’s for — daily training, contactless pay” and “Buy — €300, in stock in your country” both land, so the answer can name you and cite those facts.

Wrong slice

Only “Returns — 30 days, dispatch 1–2 days” enters the prompt. The fit and buy facts were never retrieved, and your URL is cited anyway.

A chunk is a slice of your page that an index stores and retrieves separately, rather than the whole HTML document. Your returns policy was one chunk, and “who it’s for” was another.

Before a page joins a search or RAG index, it usually gets split into smaller pieces: a title block, a spec table, a shipping accordion, a review widget, an FAQ buried at the bottom. Each piece is embedded and ranked on its own.

At query time, the system pulls the few chunks that look closest in meaning to the question. Those fragments are what get pasted into the prompt, not your carefully structured full page. This is how retrieval typically works, as the IBM source below describes it. It is not a rate this site counted.

You don’t choose the chunk size or the ranker. You can still shape which slices exist and what each one says. IBM describes the trade-off: chunks that are too large go vague, and chunks that are too small lose their meaning. Either way, the wrong slice can win the similarity match.

Same URL, two retrieval outcomes

Imagine one product page with four sections:

  1. Hero — brand, product name, one-line fit
  2. Who it’s for — daily training, contactless pay
  3. Buy — €300, stock, shops in your country
  4. Shipping and returns

A question about “a trail watch for daily training, buy in my country” should retrieve sections 2 and 3.

What often happens instead:

Outcome What entered the prompt What the answer sounds like
Fit and buy facts retrieved “Who it’s for” + “Buy” blocks Names the watch, mentions shops in your country, cites your URL for those facts
Only boilerplate retrieved Shipping/returns block only Cites your URL while discussing 30-day returns, and never names you as the pick

The crawler had the same permission and the schema was unchanged; a different slice of the page still produced a different answer.

Again, this isn’t cited versus recommended, because here your URL is the citation. What failed is which text from that URL the retriever thought was relevant.

Why the wrong slice wins

Retrieval matches on similarity. Nothing reads the whole page the way a person would.

The useful chunk usually loses for one of these reasons:

The goal is to make the facts you already have findable as their own slice.

Places and travel, same mechanic

Illustrative, not a logged test. A hotel page gets cited for “wifi and power outlets,” because the review-summary chunk matched a remote-work question, while the hours and address sat in a Maps embed the index never stored as text. Your URL is cited and your hotel is described wrongly.

A coastal walking-tour page gets linked for its scenic-route prose while the seasonal closure notice lives in a PDF linked from the footer. The answer says “open” because the retrieved chunk never saw the closure.

Commerce, places, and experiences all work the same way here, because retrieval happens at the fragment level rather than the page level.

What helps (without controlling the chunker)

  1. One scannable block per job. Put “who it’s for,” constraints, price, shop links, or seasonal hours in a single visible section, with a heading that mirrors how people actually ask. The first sentences in that block still have to answer the question if the paragraph above them never arrives, because the index stores the slice, not the story that led up to it. The same facts in a short table are often the slice that gets used.
  2. Keep the buy facts in the HTML. When buy links and stock exist only inside JavaScript cart widgets, many indexes may skip those blocks. Static HTML or JSON-LD for the same facts still matters. See your product lives in three places.
  3. Trim noisy boilerplate on money pages. Long returns policies belong on policy URLs. Repeating them high up on every product page means competing with your own product facts in the index.
  4. Check retrieval before you buy a rewrite. If the wrong fragment keeps winning, adding statistics to a blog post won’t move product facts that never entered the prompt. Fix the slice first, then apply GEO-style edits to pages that are already being retrieved.

You can’t audit every vector in every vendor’s index. You can read your own page as chunks: scroll it, imagine scissors every few paragraphs, and ask which slice would answer a real question about the watch.

If the citation points at you but the paragraph is about returns, the miss happened at the fragment level. Put the price and who it is for in a slice the index can store on its own.

Read next if: The GEO paper counted citations inside sourced answers · Your page can be used as a source without you being recommended