Articles — Can AI read your site
Can AI read your site
ChatGPT can answer without opening your page
A search-backed answer looks things up first and then writes, so it can only use text that was retrieved for that question. What the model already remembers is not the same as your live page.
You fixed the product page. The price is current, and crawlers can reach it.
Illustrative example, not a logged test. Someone asks ChatGPT where to buy a trail watch under €300 in your country. A famous brand comes back. Northloop, a fictional watch brand on this site, does not.
The page was reachable the whole time. This particular answer simply never pulled it in, because it worked from training memory, or from pages someone else had retrieved.
So this is not the access problem, where bots are blocked. It is also not cited versus named: a citation is your page used as a source, and being named is your product appearing as a pick. That split is a different miss. This one happens at the retrieve step, because a search-backed answer can only use text that was looked up for this question.
That lookup-then-write loop is what people mean by RAG (retrieval-augmented generation). You do not need the acronym to use the idea.
What RAG means here
The system searches an index, crawl results, or a shopping catalog. Only a few page slices survive the ranker for this question.
The question and the retrieved snippets get pasted together. Your whole site is not in there, only what was looked up for this answer.
The model writes the reply from that combined context. If your shop never entered the prompt, GEO rewrites and schema fixes had nothing to act on here.
RAG stands for retrieval-augmented generation, and it describes a simple loop: look up relevant text from an outside source, add it to the prompt, then write the answer.
It does not mean the model remembers your site from training, and it does not mean your JSON-LD flows in by itself.
IBM’s maintained explainer on retrieval-augmented generation walks through the same loop, and the term itself comes from Lewis et al., 2020.
Each of the big assistants wires this up differently, so a behavior you see repeatedly is still not evidence of one shared algorithm inside those companies.
Training memory is not your live page
A foundation model knows patterns up to its training cut-off: popular brands, generic advice, prices that were public at the time. That memory fills gaps whenever nothing gets retrieved.
It can also make an answer confidently wrong, naming a discontinued SKU, a shop that stopped stocking the line, or ignoring a seasonal closure it never re-checked.
When browsing, search, or a shopping tool is switched on, the assistant is supposed to lean on fresh retrieval first. Your live product page, your opening hours, or your merchant offer only count if they land in the retrieved set for that specific question.
When they don’t, GEO rewrites, schema fixes, and llms.txt maps have nothing to act on in that answer, because none of your text was in the prompt. Schema here means structured facts on the page, often a JSON-LD block that states price and identity for machines. A feed is a catalog file of price and stock. None of those enter the answer unless this question retrieved them. When browsing did run, a reported case shows the same loop from the other side. ChatGPT named a fake deodorant, Morrowen, from that brand’s own pages.
Where this sits in the stack
| Layer | Question it answers |
|---|---|
| Access | Can the system fetch or index this URL at all? |
| Retrieval (this article) | Did relevant text from you enter this prompt? |
| Cited vs named | Was your URL a source, or your brand on the shortlist? |
| GEO-style edits | If you were already retrieved, can the model lean on you more? |
Check in this order. Fix access, then ask whether retrieval happened at all, and only then worry about citation share or getting named. How AI chatbots build a shortlist covers the whole pattern. Retrieval bias, in that article, means the assistant can only recommend from what its tools actually surfaced.
Three quick examples
Illustrative example, not a logged test. Northloop is a fictional example brand on this site.
Trail watch. Northloop’s page has correct JSON-LD and lets search bots in. A ChatGPT answer with search off recites a famous brand from memory. The same question in Perplexity pulls fresh pages, so Northloop might appear, or might lose to whatever the index surfaced. Different retrieval, different set of text in the prompt, different answer.
Old Town walking tour. The municipal page lists March opening hours in plain HTML. An assistant with no live fetch repeats “open year-round” from an old guide in its training data. Retrieval would have fixed that, and the lack of retrieval is exactly why it didn’t.
Wine deal check. Your shop’s €14 offer is right there on the live product page. The model quotes €19 because it never retrieved this week’s price and fell back on a typical retail range from memory. Your feed row was perfect and had no effect on that answer.
What to check before you rewrite copy
- Was this answer search-backed? Citations, a visible “searching…” step, or facts that couldn’t plausibly come from memory all suggest retrieval ran. A fluent answer with no sources may well be memory only.
- Are you in the index path at all? If retrieval bots are blocked, or your HTML comes back empty to crawlers, stop here and read the access article.
- If you were retrieved, was it the right part of the page? Whole pages rarely enter intact, because indexes store fragments. See AI often reads only part of your page.
- Keep the measurements separate. Getting retrieved helps sourced answers and can feed shortlists, but it doesn’t guarantee you get named. Cited vs recommended is still a different scoreboard.
Before you pay for another “optimize for AI search” rewrite, answer a simpler question first: did that answer even involve a live fetch that could have included you?
Read next if: AI often reads only part of your page · Your catalog can be perfect and still never show up