Glossary
Glossary
Terms as this site uses them, grouped like the articles. 81 terms in 10 groups.
Named / recommended
You are one of the picks the assistant names on the shortlist — the product, shop, or place someone might actually choose.
Cited / linked
Your page supports the answer as a source, but the thing you sell never reaches that list.
Your page can be used as a source without you being recommended walks through the distinction on each surface.
01
How shortlists work 20 terms
Why these answers matter, whether you can trust them, and how AI chooses what to recommend.
How shortlists work 20 terms
In the reading list: How shortlists work — start with What I actually ask AI
-
AI chatbot
A tool you type a question into — ChatGPT, Claude, Perplexity, Gemini, Grok — that answers in sentences and usually hands back a few recommendations rather than a page of blue links. This site is about those recommendations: how they got there, and whether the facts hold up.
-
Assistant
The same kind of product as an AI chatbot — ChatGPT, Claude, Perplexity, Gemini, Grok, or another like them. “The assistant” means whichever one answered you in that conversation: the one that named a place, quoted a price, or said a shop was open. That is the product you used, not the LLM model inside it and not the chat window.
-
Shortlist
What an assistant actually hands back when you ask a real question: the small set of things (products, shops, places, options) it names. Everything else on this site is about how to end up on that list, with the facts right.
-
AI visibility
Showing up in AI answers with the facts right — price, shop, stock, hours, location, identity. For people who ask, that is whether the answer can be trusted. For shops and places, that is two jobs: being named, and being described correctly. Most vendor dashboards only measure the first half, and they do it badly.
-
Named / recommended
Being one of the actual picks the assistant names — the product, shop, or place someone might choose — rather than merely cited or linked, where a page is used as a source but the thing it’s about never makes the list. This is the most important split on the site.
See: Your page can be used as a source without you being recommended, Those AI visibility tools just count mentions
-
Cited / linked
A page or source used to build an answer, as a footnote or supporting link, without the brand, place, or product being named as a pick. Useful for guides; not the same prize as a shortlist slot.
See: Your page can be used as a source without you being recommended
-
Entity
One real thing the assistant can name: this watch, this hotel, this bottle at this shop. A keyword is only the text that matches. Systems have to decide whether scattered mentions describe one entity or several before they can recommend anything with confidence.
-
Keyword
A text string that matches content — “trail watch,” “creatine,” “open late” — without pinning down one specific thing. The assistant still has to know which product, shop, or address that string belongs to.
-
Knowledge graph
The matching job systems run in the background, assembling links between names, places, products, and sources from your site, Maps, feeds, and reviews. Your job is to publish one consistent story about each thing so those links can hold.
-
Corroboration
Whether enough independent, trusted sources agree on the same thing. One of the main inputs to a shortlist — famous options often win on corroboration alone, not merit — and an input that can be faked (see astroturfing, under Planted sources). It can also fail the other way: first-party copy treated as if independent sources had already agreed. For what you can change without planting: Reviews help only on the site that chatbot already reads.
See: How AI chatbots build a shortlist, Reviews help only on the site that chatbot already reads
-
Constraint fit
Whether an option actually matches what was asked — “under €300,” “shops in your city,” “not the tourist loop,” “coffee on this road,” or a symptom-specific need (“sensitive skin, no baking soda”) versus a broad “best X.” Options that don’t clearly satisfy a stated constraint tend to get dropped rather than guessed at.
See: How AI chatbots build a shortlist, “Best things to do near me” is how you get the tourist list, Ask for a white t-shirt and you get Uniqlo, Say what you want the answer to do
-
Entity clarity
How obviously “one thing” something is across the sources a system can see. The same brand spelled three ways, or a hotel listed at two different addresses, makes a system hesitant even when everything else checks out.
-
Fresh facts / described correctly
Whether the system can state accurate, current details when it names you — price, stock, hours, address, season, which shop, ship-to, returns, and where to book or buy. Stale or wrong facts destroy trust even when you make the list; missing next-step facts can mean you are recommended and still not chosen.
-
Google Business Profile / Maps listing
Google’s record for a place — name, pin, hours, reviews — that feeds Maps and many local shortlists. Thin, wrong, or inconsistent listings are a common “skipped or misdescribed” failure for restaurants, hotels, shops, and other local businesses.
-
Retrieval bias
Assistants recommend from what their tools actually surfaced — search results, Maps, shopping feeds, forum threads — not from some neutral picture of reality. Being genuinely better than the competition doesn’t help if you were never in the retrieved set to begin with.
-
Source weighting
Whether a retrieved source is treated as trustworthy and relevant for the question — or merely present. Failures include a joke treated as cooking advice, a bad packing list treated as a route, and first-party marketing copy restated as a vetted product pick.
-
Generative engine
A system that writes an answer from looked-up text instead of returning a ranked list of links. Examples include ChatGPT, Perplexity, Google’s AI Overviews, and Gemini.
-
RLHF (reinforcement learning from human feedback)
After the base model is trained, this step rewards answers people prefer. That is one reason assistants would rather sound sure than say they don’t know. The training often uses a reward model scored on comparison data.
-
Sycophancy
When the assistant agrees with you, or gives a convincing wrong answer you seem to want, instead of correcting you. Documented in models tuned with RLHF; distinct from a simple factual error.
-
Abstention
Saying “I don’t know” instead of guessing. Scoring and tuning often punish that and reward a fluent wrong answer instead.
02
Whether that list is for you 1 term
When you did not say who you are, the assistant often fills the gap with a generic profile. The term below is Default person.
Whether that list is for you 1 term
In the reading list: Whether that list is for you — start with "Best things to do near me" is how you get the tourist list
-
Default person
When you did not state age, budget, country, mobility, or trip stage, the assistant still has to pick options — and it often fills the gap with a generic traveler or buyer profile. The answer can sound personal while it was written for a generic person.
03
Can AI read your site 16 terms
Even a good reputation fails if ChatGPT never opened your page.
Can AI read your site 16 terms
In the reading list: Can AI read your site — start with ChatGPT can answer without opening your page — try the free fetch check
-
Training crawl
A bot visiting your site to gather content that may be used to train or fine-tune a model. Distinct from a retrieval / search index crawl (content that can surface in live, sourced answers) and a user-triggered fetch (a bot visiting one URL because a person asked the assistant to open it). Same company, often three different bots, three different jobs — blocking one doesn’t mean you’ve blocked the others.
-
User-triggered fetch
A bot visit that happens because someone, in that moment, asked the assistant to open a specific URL. OpenAI’s ChatGPT-User and Perplexity’s Perplexity-User are examples; both providers say robots.txt may not apply to these fetches.
-
AI crawler user-agents
The named bots (or tokens) that identify each crawl job. OpenAI splits GPTBot (training), OAI-SearchBot (index), and ChatGPT-User (user-triggered fetch); Google uses the Google-Extended robots token plus Googlebot; Anthropic lists ClaudeBot, Claude-SearchBot, and Claude-User; Microsoft’s Bingbot indexes Bing Search and Copilot-related surfaces; Perplexity uses PerplexityBot and Perplexity-User. New names appear regularly — check official docs. Last checked: September 2026.
Company Training Search / retrieval User-triggered OpenAI GPTBot OAI-SearchBot ChatGPT-User Google Google-Extended (token) Googlebot — Anthropic ClaudeBot Claude-SearchBot Claude-User Microsoft — Bingbot — Perplexity — PerplexityBot Perplexity-User Common Crawl CCBot — — Google-Extended is a robots.txt product token, not a separate HTTP user-agent — Googlebot still does the fetching. Disallowing it opts out of some Gemini training and grounding uses; it does not remove you from Google Search or AI Overviews. Bingbot is a mixed-use crawler for Bing Search and Copilot-related indexing; treat it like other search bots when you split training from retrieval in robots.txt.
See: Your catalog can be perfect and still never show up. Check how your robots.txt treats these bots
-
robots.txt
A public file that asks each crawler which pages it may fetch. Declared policy and what the server actually serves can diverge, so check the logs as well as the file.
See: Your catalog can be perfect and still never show up. Run the fetch check
-
llms.txt
A proposed (not official, not required) convention — a curated Markdown map at
/llms.txtpointing an already-permitted agent to a site’s cleanest, most useful pages. It has nothing to do with whether a crawler is allowed in; it only helps once access is already granted. llmstxt.org.See: Your catalog can be perfect and still never show up. See if yours is present
-
WAF / bot fight
Firewall or CDN rules that block automated clients while humans get through. Server logs may show 403 for bot user-agents and 200 for browsers — perfect schema behind “bot fight” mode is still invisible to that crawler.
-
Mixed-use crawler
A single bot identity used for more than one job — search indexing and training together, for example. On Cloudflare, starting 15 September 2026, new defaults apply the most restrictive applicable rule to mixed Search+Training crawlers (including Googlebot and Bingbot): blocking Training can block the whole bot unless you opt out or the provider splits by function. Cloudflare-specific; not a web-wide law.
-
Pay-Per-Crawl
Cloudflare’s closed-beta scheme (announced 1 July 2025) letting site owners charge AI crawlers per fetch via HTTP 402 / payment headers. Background for the access layer — not something most operators configure today. Distinct from the July 2026 Search / Agent / Training traffic controls on the same platform.
-
Crawl-to-refer ratio
How many pages a company’s bots take compared with how many visitors they send back. Cloudflare Radar publishes industry baselines; the number shifts with the date window and with in-app clicks that omit a Referer header. A training bot’s ratio is not the same claim as “live shopping-agent traffic is additive.”
-
AI Labyrinth
Cloudflare opt-in defense: invisible decoy links lead non-compliant crawlers into AI-generated maze pages humans never see — bot detection, not a substitute for robots.txt. Cloudflare blog.
-
Content licensing (publisher ↔ AI lab)
A commercial contract or feed that delivers journalism to an AI product outside the crawl path — so blocking GPTBot in robots.txt does not stop licensed material from reaching ChatGPT. Distinct from open-web retrieval.
-
RAG (retrieval-augmented generation)
The loop behind most search-backed AI answers: retrieve relevant text from somewhere external, add it to the prompt, then generate the reply from that combined context. The term comes from Lewis et al., 2020. Your page being reachable and accurate doesn’t matter if it was never actually retrieved for that specific question.
-
Grounded / ungrounded
An answer is grounded when it’s built from text the system actually looked up for that question — you’ll often see citations, or a “searching…” indicator. An ungrounded answer comes from the model’s training memory alone, which can be fluent, confident, and stale all at once. Neither this site nor most public documentation gives you a reliable way to tell which mode produced a given answer from the outside; treat a fluent answer with zero sources as a signal, not proof.
-
Chunk / chunking
Assistants rarely read your whole page. They store and look up smaller pieces, so a citation can point at you while the useful sentence never entered the answer. Those pieces are often stored as embeddings (numeric representations used for similarity search).
-
Foundation model / training cutoff
The base model an assistant runs on, trained on a snapshot of text up to some date (the cutoff). Anything that changed after that date — a price, a closure, a discontinued model — the model can only know about if it was retrieved live, not recalled from memory.
-
Browser agent
An assistant that opens the page itself in a browser and tries to finish a task, such as a booking or a click on Save. Different from a crawler, which fetches pages, and different from an MCP tool call, which asks a connected server for labeled facts. A browser agent often arrives with an ordinary browser user-agent. Last checked against Merj’s write-up: September 2026.
04
A citation is not a recommendation 8 terms
Your URL can show up as a source while the thing you sell is never named.
A citation is not a recommendation 8 terms
In the reading list: A citation is not a recommendation — start with Your page can be used as a source without you being recommended
-
GEO (Generative Engine Optimization)
Changing how a page is written and structured so that when a generative engine builds a sourced answer, that page is more likely to be used, cited, and given weight. Coined in a Princeton / Georgia Tech / IIT Delhi paper (arXiv:2311.09735, KDD 2024). Important limit: their tests started from pages already in the retrieved set. A rewrite does nothing if the assistant never opened the page.
-
AEO (Answer Engine Optimization)
A near-synonym for GEO, more often used around Perplexity-style cited answers and Google’s AI Overviews. You’ll also see LLMO for roughly the same territory. None of these three has become a settled industry standard — treat them as overlapping, not precisely distinct.
-
SEO
Classic search optimization: ranked lists of links, snippets, and click-through. GEO fights for share of a woven answer instead. The prizes differ, and the tactics overlap only where being cite-able and being findable both need clear, trustworthy pages.
-
AI Overview
Google’s generated summary block at the top of search results, with supporting links underneath. One instance of a generative engine, not a synonym for the category.
See: Your page can be used as a source without you being recommended
-
E-E-A-T
Google’s framework (Experience, Expertise, Authoritativeness, Trustworthiness) for judging whether content deserves trust in classic search. On this site, corroboration is the plain-language term for whether independent sources agree on the same thing.
-
Zero-click
A search or AI interaction that fully answers the question without the person visiting the source site. Generative answers increase this by design — being cited no longer reliably sends a visitor your way.
See: Your page can be used as a source without you being recommended
05
Planted sources 8 terms
Some of the trust AI reads was put there on purpose — including on a brand's own site.
Planted sources 8 terms
In the reading list: Planted sources — start with ChatGPT recommended a deodorant that didn’t exist
-
Retrieval poisoning
Corrupting the pool of pages a system might retrieve — so one planted snippet on a high-overlap URL (often Reddit or forum UGC) skews many related answers. Plain-language name for what the WARP paper formalizes; distinct from prompt injection in a single chat.
-
WARP (Web Agent Retrieval Poisoning)
A named attack from a May 2026 Cornell Tech paper (arXiv:2605.24245): deep-research agents reuse the same handful of retrieved UGC pages across many related queries in one session, so a short crafted snippet on one page can get cited or promoted across all of them. The bottleneck is exposure (getting retrieved), not persuasion. Tested on STORM, Co-STORM, and OmniThink; commercial Deep Research was reconnaissance only.
-
Retrieval overlap
Within one research session or topic cluster, the same URLs — often Reddit or forum pages — showing up again and again across related queries. That overlap is the structural vulnerability WARP exploits; one poisoned page can affect many answers.
-
Deep-research agent
A system that runs multiple rounds of retrieval and synthesis to produce a structured report, rather than a single-turn chat answer. STORM, Co-STORM, and OmniThink are open-source examples from the WARP paper.
-
UGC (user-generated content)
Pages and threads written by users on platforms like Reddit, forums, and Wikipedia — as opposed to editorial or merchant-owned pages. High overlap in deep-research retrieval makes UGC a concentrated attack surface.
-
Data poisoning
Corrupting the data a system learns from or looks up so its output is manipulated. Retrieval poisoning corrupts the pool of pages future answers might pull from. Prompt injection corrupts what one model sees in a single chat or document.
-
Astroturfing
Manufacturing the appearance of organic, independent agreement when it’s actually coordinated and often paid — fake grassroots enthusiasm. The mechanism behind planted Reddit threads that read as genuine discussion until a brand mention gets worked in.
-
Prompt injection
Manipulating the instructions a model sees within a single conversation or document — distinct from poisoning the broader pool of content a system might retrieve across many future conversations. WARP is closer to the latter.
See: Companies use Reddit to change what AI recommends. The assistant said the table was booked
06
Why you were skipped 5 terms
Work out why you were skipped before you rewrite a page.
Why you were skipped 5 terms
In the reading list: Why you were skipped — start with Find out why you were skipped before you rewrite the page
-
Prompt audit / AI visibility check
Running a few real buyer questions in the tools people actually use, writing down who was named, who was only linked, and what was wrong — not one mystery score from a dashboard.
-
Diagnosis
Before you rewrite a page, work out why you were missing: the bot never opened you, opened you and named someone else, named you with the wrong facts, or a better-known rival won. “We’re missing” is not a brief for a rewrite.
See: Find out why you were skipped before you rewrite the page
-
Never looked up
This specific answer never pulled text from your live site — often an ungrounded reply from training memory, or retrieval that never included you. Fix starts with access and whether lookup happened at all, not copy.
See: Find out why you were skipped before you rewrite the page
-
Looked up, not named
Your page or URL was in the mix as a source, but the product, shop, or place you sell never made the shortlist — cited rather than recommended, or outgunned on corroboration and fit.
See: Find out why you were skipped before you rewrite the page
-
Named with wrong facts
You appear on the shortlist, but price, shop, stock, hours, or spec is wrong. Fix the catalog and the identity.
See: Find out why you were skipped before you rewrite the page
07
Product data for shops 11 terms
Identity for a product, a shop, or a place, then price, stock, and identity on the webpage, in structured data, and in your feed.
Product data for shops 11 terms
In the reading list: Product data for shops — start with Wrong price, wrong shop, fake discount — check JSON-LD on a URL with the fetch check
-
Structured data
Facts on a page in a shape machines can read without guessing from layout — often JSON-LD using Schema.org types. Helps the “described correctly” job; does not by itself get you named.
See: Structured data publishes facts machines can read. Check JSON-LD on a URL
-
JSON-LD
A small block of labeled, machine-readable facts embedded in a webpage (usually inside a
<script type="application/ld+json">tag), using a shared vocabulary — most often Schema.org. It tells a machine which string is the price and which is the name, instead of guessing from a hero image and marketing copy.See: Structured data publishes facts machines can read. Field-by-field Product / Offer check
-
Schema.org
The open vocabulary most JSON-LD product markup is written in — shared types and properties like
ProductandOfferthat many systems already know how to read. -
Product feed
A structured file or API (CSV, XML, etc.) listing product records in bulk — one row per SKU — used by shopping systems like Google Merchant Center. Distinct from a webpage or JSON-LD: the bulk catalog pipe, and it needs to agree with both.
-
Offer
In Schema.org, the block that carries price, currency, availability, and condition for a product — usually nested under
Product. -
Availability
Whether an item is in stock, out of stock, or pre-order — in JSON-LD (
Offer.availability) and in feed fields. Stale availability is a common wrong-fact failure. -
SKU
Stock Keeping Unit — your internal identifier for one sellable variant. Needs to stay stable and match across webpage, schema, and feed.
-
GTIN
Global Trade Item Number — the barcode-family identifier (EAN/UPC) that says “this exact trade item,” not a similarly-named lookalike.
-
MPN
Manufacturer Part Number — pairs with brand to identify a specific product when no GTIN exists, common for private-label goods.
-
Condition
new/refurbished/usedin a product feed. Without it, a cheap used listing can quietly poison a “best new watch under €300” shortlist by looking like the same offer. -
Merchant Center
Google’s system for ingesting product feeds and powering Shopping-related surfaces, including AI shopping features that lean on feed data rather than parsing a webpage directly.
08
Shopping inside chat 6 terms
Live stock lookup and checkout in chat help after someone already picked you. They do not replace retrieval, fit, and trust on the public shortlist. The tool-calling article is the page to hand a developer. Product facts are in the previous group.
Shopping inside chat 6 terms
In the reading list: Shopping inside chat — start with ChatGPT can look up your stock only after you connect it
-
Chat
The conversation window you type into. Checkout in chat means paying without leaving that window. ChatGPT is a product; the window is chat.
-
MCP (Model Context Protocol)
An open standard (Anthropic-originated; stewarded by the Agentic AI Foundation, a Linux Foundation directed fund, from December 2025) letting an AI application call external tools and data, such as a live catalog, a database, or an API, during a conversation. It answers whether this connected assistant can look up your facts on demand. It does not answer whether ChatGPT will recommend you to a stranger.
See: ChatGPT can look up your stock only after you connect it. The assistant checks your stock by calling a tool you publish
-
WebMCP
A proposed way for a page to expose actions to a browser agent, so an agent that is already on the page can call a typed action instead of clicking around. Chromium-centric as of autumn 2026. It covers acting after arrival. It does not affect whether an assistant names you.
-
Agentic commerce
The umbrella term for AI systems acting on a person’s behalf in a shopping context. Two different jobs get sold under the same phrase: getting recommended by a third-party assistant (ChatGPT, Perplexity, Gemini), and making your own site’s search or catalog directly queryable by an agent that already arrived with intent. Under that umbrella sit live catalog lookup (MCP), checkout (ACP/UCP), on-site agent-native search, and a browser agent that uses the page itself. That browser route sits beside live lookup and checkout, and it still does not get you named. Ask which specific piece a vendor means before believing “agentic-commerce-ready” implies anything about being recommended.
See: Checkout in chat comes after the shortlist. Your catalog can be perfect and still never show up. The assistant said the table was booked
-
ACP (Agentic Commerce Protocol)
OpenAI and Stripe’s open protocol (beta) for completing a purchase inside a chat — cart, payment, order — once a shopper has already chosen a product. OpenAI has since added product-feed paths for discovery; still not the same as being named on a stranger’s shortlist.
-
UCP (Universal Commerce Protocol)
Google’s open standard for a broader shopping interaction — discovery through post-purchase — on AI Mode, Gemini, and partner surfaces. Checkout infrastructure and catalog pipes, not a recommendation mechanism.
09
Places and travel 3 terms
Restaurants, hotels, walking tours, and trips: different directories, wrong hours, and a bias toward famous places.
Places and travel 3 terms
In the reading list: Places and travel — start with Why you got the same five landmarks
-
Digital overtourism
An algorithmically driven concentration of destination visibility before anyone travels — the same iconic places named again and again in AI travel answers, which can preload demand at spots that were already crowded. Distinct from physical overtourism on the ground, but related.
-
Local directory
The place database a specific chatbot uses for local answers — Foursquare and Yelp for ChatGPT search, Yelp and Tripadvisor for Perplexity, Google Maps for many Gemini answers.
-
LocalBusiness
A Schema.org type for a physical place on your own webpage — name, address, phone, hours — usually as JSON-LD. Prefer a specific subtype (
Restaurant,Store) over the generic type. Hours and address have to match your Google Business Profile and other listings; stale markup is the same failure as stale directory data. ChatGPT and Perplexity often lean on licensed directories instead of your site markup.
10
Adjacent terms 3 terms
Words you will see elsewhere; not all have a dedicated article yet.
Adjacent terms 3 terms
Lookup only — terms you will see elsewhere; not every entry has a dedicated article yet.
-
Hallucination
When a model states something confidently and fluently that isn’t true — a wrong price, a nonexistent place, a discontinued SKU described as current. Almost every failure mode this site documents is a specific, diagnosable cause of hallucination rather than a mysterious glitch.
-
LLM (large language model)
The model inside an AI chatbot or assistant, trained on very large amounts of text so it can predict and write language. ChatGPT, Claude, Gemini, and similar products are built on LLMs.
-
AI watermarking / content provenance
Technical methods for marking AI-generated content so it can be identified as synthetic — C2PA metadata, statistical signals like SynthID, and similar. The EU AI Act’s Article 50 transparency rules apply from 2 August 2026 for providers serving EU users; marking is imperfect and often stripped in transit. Relevant background for fake media and deepfakes.
Missing something you'd expect to see here? That probably means an article assumes a term it never defines — tell us.