Articles — Can AI read your site
Can AI read your site
Your catalog can be perfect and still never show up
The name and the price on your page can be right, and you still never appear in the answer, because nothing was allowed to open or index the page.
The name, the price, and whether it is in stock all match on the page.
Illustrative example, not a logged test. Someone asks ChatGPT or Perplexity for a running watch under €300, and the Northloop Trail Watch S — a fictional example brand used across this site — never appears in the answer.
The data was right. The system may never have fetched it, or never indexed it for the tool that answered.
That is a different failure from find out why you were skipped. You can still lose on fame, identity, or stale facts even when crawlers can reach you. Below: what gets blocked, what gets indexed, and what only appears after JavaScript runs in a browser.
Perfect facts, missing fetch
Schema (structured facts on the page, often as JSON-LD) and a product feed (a catalog file of price and stock that shopping tools read) answer: “If a system reads this URL or catalog row, what should it believe?”
They do not answer: “Did any system allowed to answer shopping questions actually fetch or index this site?”
Three separate jobs show up in the wild:
| Job | Plain meaning | Example (OpenAI) |
|---|---|---|
| Training crawl | Content may enter foundation-model training | GPTBot |
| Retrieval / search index | Pages can surface in search-backed answers and citations | OAI-SearchBot |
| User-triggered fetch | A user asked the tool to open a specific URL | ChatGPT-User |
OpenAI says robots.txt may not apply to user-triggered fetches, because that visit is for this conversation rather than the search index.
A training crawl takes pages and rarely sends a visitor back. A live shopping or booking agent can look like a shopper arriving through a different door. A claim that agents are additive and a claim that they only take pages are usually describing different jobs. A token here is the bot name in robots.txt, such as GPTBot or OAI-SearchBot.
Anthropic, Google, and Perplexity use similar splits under different names, and each has its own robots.txt token. Blocking “AI” in one line often blocks the wrong thing, or fails to block the thing you think.
OpenAI’s own docs: Overview of OpenAI Crawlers. They also publish OAI-AdsBot for checking ad landing pages, which is neither the training crawl nor the search index.
Why that makes you invisible
Page, schema, and feed agree on €24, in stock. OAI-SearchBot is disallowed in robots.txt, so ChatGPT search never indexes the shop.
An official page exists, but the offer and how to book only appear after login. Crawlers never see it, so the model defaults to famous names.
GPTBot is allowed, and OAI-SearchBot is blocked by mistake. Training may see the site; search-backed answers will not cite it.
Blocked retrieval bot. You disallow OAI-SearchBot in robots.txt to keep ChatGPT out, but that token controls ChatGPT search indexing, not every OpenAI visit. Block the search bot and you opt out of one path into answers. A copied robots.txt that blocks every AI crawler can block the search bot by accident.
Training opt-out is not a visibility opt-out. Disallowing GPTBot says “don’t use this in training.” It does not automatically remove you from live search indexes, and allowing GPTBot does not guarantee you appear in answers. Treat each token on purpose. A Disallow line does not always mean the page is gone from every answer. Google’s AI features can still draw on the ordinary Search index, built by Googlebot, and a license can feed the same text through a contract instead of a crawl.
Cloudflare switches. If you use Cloudflare, Search, Agent, and Training are separate controls, not one “block AI” button (Your site, your rules). Blocking Training on a mixed-use crawler such as Googlebot can also stop that crawler from indexing you for search (Have it both ways).
WAF or firewall. A web application firewall can block automated clients while people get through. Server logs show 403 for bot user-agents while humans get 200. Perfect schema behind Cloudflare “bot fight” mode is still invisible to that crawler.
Login, heavy JavaScript, or thin HTML. The crawler gets a shell page, not the price. Most AI crawlers do not run JavaScript. SearchVIU counted crawlers in November 2025 and found about 69% could not execute it, including GPTBot, ClaudeBot, and PerplexityBot. Googlebot is an exception. A page whose price loads only in the browser is an empty page to those bots. The free Fetch check flags thin visible HTML as a hint of that failure.
Never in that tool’s index. Perplexity, Gemini, and ChatGPT do not share one retrieval system. Perfect data in Google Merchant Center does not automatically populate ChatGPT search. See how AI builds a shortlist for the working model, rather than one verified algorithm inside those companies.
robots.txt vs llms.txt
robots.txt tells automated crawlers what they may fetch. It is a permission layer: allow or disallow by user-agent. It does not describe your catalog or point agents to the right pages.
llms.txt (proposed convention, llmstxt.org, v2 updated August 2026) is a curated Markdown map at /llms.txt, with links to LLM-friendly pages, short summaries, and optional sections. It helps an agent that is already allowed to visit find the useful pages faster. It is not a substitute for robots.txt, and it is not a ranking switch.
Limits as of 2026:
- Google Search does not use llms.txt; Google’s public line is it neither helps nor hurts traditional search rankings.
- Adoption is strongest on developer docs and coding-agent sites; e-commerce is thinner.
- v2 adds discoverability hints (
rel=alternatefor Markdown twins, subpath files). Still young, useful for docs-like surfaces, optional for most shops.
If retrieval bots are blocked, llms.txt changes nothing. If they are allowed, llms.txt is an optional map rather than a visibility trick.
Official crawler docs and llms.txt are also on Resources.
What to check
- Read robots.txt. Which user-agents are disallowed? Do you mean to block training only, or search indexing too?
- Or run the free fetch check on your domain — robots.txt, llms.txt, sitemap, JSON-LD, and a thin-HTML heuristic. It does not tell you whether AI would recommend you.
- Check server logs (or ask your host): do
OAI-SearchBot,GPTBot, PerplexityBot, or equivalents get 200 on product URLs, or 403? - Fetch a product URL as a bot would (curl with the published user-agent string). Is the price and stock in the HTML, or only after JavaScript?
- Separate the fixes: allow the retrieval bots you care about, keep a training opt-out if you want one, fix WAF false positives, then worry about llms.txt or Markdown twins if you run docs-heavy content.
- Remember corroboration: even a perfectly fetched page loses to a famous listing with more reviews. Access is necessary, not sufficient.
Access is the layer under your product lives in three places. Get fetch and index paths open, then keep facts consistent everywhere else.
Read next if: ChatGPT can answer without opening your page · AI often reads only part of your page