Glossary

Glossary

Terms as this site uses them, grouped like the articles. 81 terms in 10 groups.

01

How shortlists work 20 terms

Why these answers matter, whether you can trust them, and how AI chooses what to recommend.

In the reading list: How shortlists work — start with What I actually ask AI

  • AI chatbot

    A tool you type a question into — ChatGPT, Claude, Perplexity, Gemini, Grok — that answers in sentences and usually hands back a few recommendations rather than a page of blue links. This site is about those recommendations: how they got there, and whether the facts hold up.

    See: How AI chatbots build a shortlist

  • Assistant

    The same kind of product as an AI chatbot — ChatGPT, Claude, Perplexity, Gemini, Grok, or another like them. “The assistant” means whichever one answered you in that conversation: the one that named a place, quoted a price, or said a shop was open. That is the product you used, not the LLM model inside it and not the chat window.

    See: How AI chatbots build a shortlist

  • Shortlist

    What an assistant actually hands back when you ask a real question: the small set of things (products, shops, places, options) it names. Everything else on this site is about how to end up on that list, with the facts right.

    See: How AI chatbots build a shortlist

  • AI visibility

    Showing up in AI answers with the facts right — price, shop, stock, hours, location, identity. For people who ask, that is whether the answer can be trusted. For shops and places, that is two jobs: being named, and being described correctly. Most vendor dashboards only measure the first half, and they do it badly.

    See: How AI chatbots build a shortlist

  • Named / recommended

    Being one of the actual picks the assistant names — the product, shop, or place someone might choose — rather than merely cited or linked, where a page is used as a source but the thing it’s about never makes the list. This is the most important split on the site.

    See: Your page can be used as a source without you being recommended, Those AI visibility tools just count mentions

  • Cited / linked

    A page or source used to build an answer, as a footnote or supporting link, without the brand, place, or product being named as a pick. Useful for guides; not the same prize as a shortlist slot.

    See: Your page can be used as a source without you being recommended

  • Entity

    One real thing the assistant can name: this watch, this hotel, this bottle at this shop. A keyword is only the text that matches. Systems have to decide whether scattered mentions describe one entity or several before they can recommend anything with confidence.

    See: The assistant names a product, a shop, or a place

  • Keyword

    A text string that matches content — “trail watch,” “creatine,” “open late” — without pinning down one specific thing. The assistant still has to know which product, shop, or address that string belongs to.

    See: The assistant names a product, a shop, or a place

  • Knowledge graph

    The matching job systems run in the background, assembling links between names, places, products, and sources from your site, Maps, feeds, and reviews. Your job is to publish one consistent story about each thing so those links can hold.

    See: The assistant names a product, a shop, or a place

  • Corroboration

    Whether enough independent, trusted sources agree on the same thing. One of the main inputs to a shortlist — famous options often win on corroboration alone, not merit — and an input that can be faked (see astroturfing, under Planted sources). It can also fail the other way: first-party copy treated as if independent sources had already agreed. For what you can change without planting: Reviews help only on the site that chatbot already reads.

    See: How AI chatbots build a shortlist, Reviews help only on the site that chatbot already reads

  • Constraint fit

    Whether an option actually matches what was asked — “under €300,” “shops in your city,” “not the tourist loop,” “coffee on this road,” or a symptom-specific need (“sensitive skin, no baking soda”) versus a broad “best X.” Options that don’t clearly satisfy a stated constraint tend to get dropped rather than guessed at.

    See: How AI chatbots build a shortlist, “Best things to do near me” is how you get the tourist list, Ask for a white t-shirt and you get Uniqlo, Say what you want the answer to do

  • Entity clarity

    How obviously “one thing” something is across the sources a system can see. The same brand spelled three ways, or a hotel listed at two different addresses, makes a system hesitant even when everything else checks out.

    See: The assistant names a product, a shop, or a place

  • Fresh facts / described correctly

    Whether the system can state accurate, current details when it names you — price, stock, hours, address, season, which shop, ship-to, returns, and where to book or buy. Stale or wrong facts destroy trust even when you make the list; missing next-step facts can mean you are recommended and still not chosen.

    See: You can be recommended and still not get chosen

  • Google Business Profile / Maps listing

    Google’s record for a place — name, pin, hours, reviews — that feeds Maps and many local shortlists. Thin, wrong, or inconsistent listings are a common “skipped or misdescribed” failure for restaurants, hotels, shops, and other local businesses.

    See: AI place answers come from different directories

  • Retrieval bias

    Assistants recommend from what their tools actually surfaced — search results, Maps, shopping feeds, forum threads — not from some neutral picture of reality. Being genuinely better than the competition doesn’t help if you were never in the retrieved set to begin with.

    See: How AI chatbots build a shortlist

  • Source weighting

    Whether a retrieved source is treated as trustworthy and relevant for the question — or merely present. Failures include a joke treated as cooking advice, a bad packing list treated as a route, and first-party marketing copy restated as a vetted product pick.

    See: AI can give stupid or dangerous advice

  • Generative engine

    A system that writes an answer from looked-up text instead of returning a ranked list of links. Examples include ChatGPT, Perplexity, Google’s AI Overviews, and Gemini.

    See: How AI chatbots build a shortlist

  • RLHF (reinforcement learning from human feedback)

    After the base model is trained, this step rewards answers people prefer. That is one reason assistants would rather sound sure than say they don’t know. The training often uses a reward model scored on comparison data.

    See: Why AI won’t just say it doesn’t know

  • Sycophancy

    When the assistant agrees with you, or gives a convincing wrong answer you seem to want, instead of correcting you. Documented in models tuned with RLHF; distinct from a simple factual error.

    See: Why AI won’t just say it doesn’t know

  • Abstention

    Saying “I don’t know” instead of guessing. Scoring and tuning often punish that and reward a fluent wrong answer instead.

    See: Why AI won’t just say it doesn’t know

02

Whether that list is for you 1 term

When you did not say who you are, the assistant often fills the gap with a generic profile. The term below is Default person.

In the reading list: Whether that list is for you — start with "Best things to do near me" is how you get the tourist list

  • Default person

    When you did not state age, budget, country, mobility, or trip stage, the assistant still has to pick options — and it often fills the gap with a generic traveler or buyer profile. The answer can sound personal while it was written for a generic person.

    See: Without your context, the answer is for someone else

03

Can AI read your site 16 terms

Even a good reputation fails if ChatGPT never opened your page.

In the reading list: Can AI read your site — start with ChatGPT can answer without opening your page — try the free fetch check

  • Training crawl

    A bot visiting your site to gather content that may be used to train or fine-tune a model. Distinct from a retrieval / search index crawl (content that can surface in live, sourced answers) and a user-triggered fetch (a bot visiting one URL because a person asked the assistant to open it). Same company, often three different bots, three different jobs — blocking one doesn’t mean you’ve blocked the others.

    See: Your catalog can be perfect and still never show up

  • User-triggered fetch

    A bot visit that happens because someone, in that moment, asked the assistant to open a specific URL. OpenAI’s ChatGPT-User and Perplexity’s Perplexity-User are examples; both providers say robots.txt may not apply to these fetches.

    See: Your catalog can be perfect and still never show up

  • AI crawler user-agents

    The named bots (or tokens) that identify each crawl job. OpenAI splits GPTBot (training), OAI-SearchBot (index), and ChatGPT-User (user-triggered fetch); Google uses the Google-Extended robots token plus Googlebot; Anthropic lists ClaudeBot, Claude-SearchBot, and Claude-User; Microsoft’s Bingbot indexes Bing Search and Copilot-related surfaces; Perplexity uses PerplexityBot and Perplexity-User. New names appear regularly — check official docs. Last checked: September 2026.

    Company Training Search / retrieval User-triggered
    OpenAI GPTBot OAI-SearchBot ChatGPT-User
    Google Google-Extended (token) Googlebot —
    Anthropic ClaudeBot Claude-SearchBot Claude-User
    Microsoft — Bingbot —
    Perplexity — PerplexityBot Perplexity-User
    Common Crawl CCBot — —

    Google-Extended is a robots.txt product token, not a separate HTTP user-agent — Googlebot still does the fetching. Disallowing it opts out of some Gemini training and grounding uses; it does not remove you from Google Search or AI Overviews. Bingbot is a mixed-use crawler for Bing Search and Copilot-related indexing; treat it like other search bots when you split training from retrieval in robots.txt.

    See: Your catalog can be perfect and still never show up. Check how your robots.txt treats these bots

  • robots.txt

    A public file that asks each crawler which pages it may fetch. Declared policy and what the server actually serves can diverge, so check the logs as well as the file.

    See: Your catalog can be perfect and still never show up. Run the fetch check

  • llms.txt

    A proposed (not official, not required) convention — a curated Markdown map at /llms.txt pointing an already-permitted agent to a site’s cleanest, most useful pages. It has nothing to do with whether a crawler is allowed in; it only helps once access is already granted. llmstxt.org.

    See: Your catalog can be perfect and still never show up. See if yours is present

  • WAF / bot fight

    Firewall or CDN rules that block automated clients while humans get through. Server logs may show 403 for bot user-agents and 200 for browsers — perfect schema behind “bot fight” mode is still invisible to that crawler.

    See: Your catalog can be perfect and still never show up

  • Mixed-use crawler

    A single bot identity used for more than one job — search indexing and training together, for example. On Cloudflare, starting 15 September 2026, new defaults apply the most restrictive applicable rule to mixed Search+Training crawlers (including Googlebot and Bingbot): blocking Training can block the whole bot unless you opt out or the provider splits by function. Cloudflare-specific; not a web-wide law.

    See: Your catalog can be perfect and still never show up

  • Pay-Per-Crawl

    Cloudflare’s closed-beta scheme (announced 1 July 2025) letting site owners charge AI crawlers per fetch via HTTP 402 / payment headers. Background for the access layer — not something most operators configure today. Distinct from the July 2026 Search / Agent / Training traffic controls on the same platform.

  • Crawl-to-refer ratio

    How many pages a company’s bots take compared with how many visitors they send back. Cloudflare Radar publishes industry baselines; the number shifts with the date window and with in-app clicks that omit a Referer header. A training bot’s ratio is not the same claim as “live shopping-agent traffic is additive.”

  • AI Labyrinth

    Cloudflare opt-in defense: invisible decoy links lead non-compliant crawlers into AI-generated maze pages humans never see — bot detection, not a substitute for robots.txt. Cloudflare blog.

  • Content licensing (publisher ↔ AI lab)

    A commercial contract or feed that delivers journalism to an AI product outside the crawl path — so blocking GPTBot in robots.txt does not stop licensed material from reaching ChatGPT. Distinct from open-web retrieval.

    See: Your catalog can be perfect and still never show up

  • RAG (retrieval-augmented generation)

    The loop behind most search-backed AI answers: retrieve relevant text from somewhere external, add it to the prompt, then generate the reply from that combined context. The term comes from Lewis et al., 2020. Your page being reachable and accurate doesn’t matter if it was never actually retrieved for that specific question.

    See: ChatGPT can answer without opening your page

  • Grounded / ungrounded

    An answer is grounded when it’s built from text the system actually looked up for that question — you’ll often see citations, or a “searching…” indicator. An ungrounded answer comes from the model’s training memory alone, which can be fluent, confident, and stale all at once. Neither this site nor most public documentation gives you a reliable way to tell which mode produced a given answer from the outside; treat a fluent answer with zero sources as a signal, not proof.

    See: ChatGPT can answer without opening your page

  • Chunk / chunking

    Assistants rarely read your whole page. They store and look up smaller pieces, so a citation can point at you while the useful sentence never entered the answer. Those pieces are often stored as embeddings (numeric representations used for similarity search).

    See: AI often reads only part of your page

  • Foundation model / training cutoff

    The base model an assistant runs on, trained on a snapshot of text up to some date (the cutoff). Anything that changed after that date — a price, a closure, a discontinued model — the model can only know about if it was retrieved live, not recalled from memory.

    See: ChatGPT can answer without opening your page

  • Browser agent

    An assistant that opens the page itself in a browser and tries to finish a task, such as a booking or a click on Save. Different from a crawler, which fetches pages, and different from an MCP tool call, which asks a connected server for labeled facts. A browser agent often arrives with an ordinary browser user-agent. Last checked against Merj’s write-up: September 2026.

    See: The assistant said the table was booked

04

Your URL can show up as a source while the thing you sell is never named.

In the reading list: A citation is not a recommendation — start with Your page can be used as a source without you being recommended

  • GEO (Generative Engine Optimization)

    Changing how a page is written and structured so that when a generative engine builds a sourced answer, that page is more likely to be used, cited, and given weight. Coined in a Princeton / Georgia Tech / IIT Delhi paper (arXiv:2311.09735, KDD 2024). Important limit: their tests started from pages already in the retrieved set. A rewrite does nothing if the assistant never opened the page.

    See: The GEO paper counted citations inside sourced answers

  • AEO (Answer Engine Optimization)

    A near-synonym for GEO, more often used around Perplexity-style cited answers and Google’s AI Overviews. You’ll also see LLMO for roughly the same territory. None of these three has become a settled industry standard — treat them as overlapping, not precisely distinct.

    See: The GEO paper counted citations inside sourced answers

  • Share of the answer / citation share

    In GEO research, how much of a synthesized answer leans on your page among sources that were already competing in the same response. It is a share of that written answer. It does not promise traffic, and it does not mean ChatGPT will name your product.

    See: The GEO paper counted citations inside sourced answers

  • SEO

    Classic search optimization: ranked lists of links, snippets, and click-through. GEO fights for share of a woven answer instead. The prizes differ, and the tactics overlap only where being cite-able and being findable both need clear, trustworthy pages.

    See: The GEO paper counted citations inside sourced answers

  • AI Overview

    Google’s generated summary block at the top of search results, with supporting links underneath. One instance of a generative engine, not a synonym for the category.

    See: Your page can be used as a source without you being recommended

  • E-E-A-T

    Google’s framework (Experience, Expertise, Authoritativeness, Trustworthiness) for judging whether content deserves trust in classic search. On this site, corroboration is the plain-language term for whether independent sources agree on the same thing.

  • Share of voice / mention tracker

    What most paid “AI visibility” dashboards sell: send the same questions repeatedly, scan replies for a brand name, chart it over time. A logbook of those runs is useful. A single score is misleading, because it typically can’t distinguish named from merely linked, or a correct mention from wrong facts.

    See: Those AI visibility tools just count mentions

  • Zero-click

    A search or AI interaction that fully answers the question without the person visiting the source site. Generative answers increase this by design — being cited no longer reliably sends a visitor your way.

    See: Your page can be used as a source without you being recommended

05

Planted sources 8 terms

Some of the trust AI reads was put there on purpose — including on a brand's own site.

In the reading list: Planted sources — start with ChatGPT recommended a deodorant that didn’t exist

  • Retrieval poisoning

    Corrupting the pool of pages a system might retrieve — so one planted snippet on a high-overlap URL (often Reddit or forum UGC) skews many related answers. Plain-language name for what the WARP paper formalizes; distinct from prompt injection in a single chat.

    See: Companies use Reddit to change what AI recommends

  • WARP (Web Agent Retrieval Poisoning)

    A named attack from a May 2026 Cornell Tech paper (arXiv:2605.24245): deep-research agents reuse the same handful of retrieved UGC pages across many related queries in one session, so a short crafted snippet on one page can get cited or promoted across all of them. The bottleneck is exposure (getting retrieved), not persuasion. Tested on STORM, Co-STORM, and OmniThink; commercial Deep Research was reconnaissance only.

    See: Companies use Reddit to change what AI recommends

  • Retrieval overlap

    Within one research session or topic cluster, the same URLs — often Reddit or forum pages — showing up again and again across related queries. That overlap is the structural vulnerability WARP exploits; one poisoned page can affect many answers.

    See: Companies use Reddit to change what AI recommends

  • Deep-research agent

    A system that runs multiple rounds of retrieval and synthesis to produce a structured report, rather than a single-turn chat answer. STORM, Co-STORM, and OmniThink are open-source examples from the WARP paper.

    See: Companies use Reddit to change what AI recommends

  • UGC (user-generated content)

    Pages and threads written by users on platforms like Reddit, forums, and Wikipedia — as opposed to editorial or merchant-owned pages. High overlap in deep-research retrieval makes UGC a concentrated attack surface.

    See: Companies use Reddit to change what AI recommends

  • Data poisoning

    Corrupting the data a system learns from or looks up so its output is manipulated. Retrieval poisoning corrupts the pool of pages future answers might pull from. Prompt injection corrupts what one model sees in a single chat or document.

    See: Companies use Reddit to change what AI recommends

  • Astroturfing

    Manufacturing the appearance of organic, independent agreement when it’s actually coordinated and often paid — fake grassroots enthusiasm. The mechanism behind planted Reddit threads that read as genuine discussion until a brand mention gets worked in.

    See: Companies use Reddit to change what AI recommends

  • Prompt injection

    Manipulating the instructions a model sees within a single conversation or document — distinct from poisoning the broader pool of content a system might retrieve across many future conversations. WARP is closer to the latter.

    See: Companies use Reddit to change what AI recommends. The assistant said the table was booked

06

Why you were skipped 5 terms

Work out why you were skipped before you rewrite a page.

In the reading list: Why you were skipped — start with Find out why you were skipped before you rewrite the page

07

Product data for shops 11 terms

Identity for a product, a shop, or a place, then price, stock, and identity on the webpage, in structured data, and in your feed.

In the reading list: Product data for shops — start with Wrong price, wrong shop, fake discount — check JSON-LD on a URL with the fetch check

  • Structured data

    Facts on a page in a shape machines can read without guessing from layout — often JSON-LD using Schema.org types. Helps the “described correctly” job; does not by itself get you named.

    See: Structured data publishes facts machines can read. Check JSON-LD on a URL

  • JSON-LD

    A small block of labeled, machine-readable facts embedded in a webpage (usually inside a <script type="application/ld+json"> tag), using a shared vocabulary — most often Schema.org. It tells a machine which string is the price and which is the name, instead of guessing from a hero image and marketing copy.

    See: Structured data publishes facts machines can read. Field-by-field Product / Offer check

  • Schema.org

    The open vocabulary most JSON-LD product markup is written in — shared types and properties like Product and Offer that many systems already know how to read.

    See: Structured data publishes facts machines can read

  • Product feed

    A structured file or API (CSV, XML, etc.) listing product records in bulk — one row per SKU — used by shopping systems like Google Merchant Center. Distinct from a webpage or JSON-LD: the bulk catalog pipe, and it needs to agree with both.

    See: Your product lives in three places

  • Offer

    In Schema.org, the block that carries price, currency, availability, and condition for a product — usually nested under Product.

    See: Your product lives in three places

  • Availability

    Whether an item is in stock, out of stock, or pre-order — in JSON-LD (Offer.availability) and in feed fields. Stale availability is a common wrong-fact failure.

    See: Your product lives in three places

  • SKU

    Stock Keeping Unit — your internal identifier for one sellable variant. Needs to stay stable and match across webpage, schema, and feed.

    See: Your product lives in three places

  • GTIN

    Global Trade Item Number — the barcode-family identifier (EAN/UPC) that says “this exact trade item,” not a similarly-named lookalike.

    See: Your product lives in three places

  • MPN

    Manufacturer Part Number — pairs with brand to identify a specific product when no GTIN exists, common for private-label goods.

    See: Your product lives in three places

  • Condition

    new / refurbished / used in a product feed. Without it, a cheap used listing can quietly poison a “best new watch under €300” shortlist by looking like the same offer.

    See: Your product lives in three places

  • Merchant Center

    Google’s system for ingesting product feeds and powering Shopping-related surfaces, including AI shopping features that lean on feed data rather than parsing a webpage directly.

    See: Your product lives in three places

08

Shopping inside chat 6 terms

Live stock lookup and checkout in chat help after someone already picked you. They do not replace retrieval, fit, and trust on the public shortlist. The tool-calling article is the page to hand a developer. Product facts are in the previous group.

In the reading list: Shopping inside chat — start with ChatGPT can look up your stock only after you connect it

  • Chat

    The conversation window you type into. Checkout in chat means paying without leaving that window. ChatGPT is a product; the window is chat.

    See: Checkout in chat comes after the shortlist

  • MCP (Model Context Protocol)

    An open standard (Anthropic-originated; stewarded by the Agentic AI Foundation, a Linux Foundation directed fund, from December 2025) letting an AI application call external tools and data, such as a live catalog, a database, or an API, during a conversation. It answers whether this connected assistant can look up your facts on demand. It does not answer whether ChatGPT will recommend you to a stranger.

    See: ChatGPT can look up your stock only after you connect it. The assistant checks your stock by calling a tool you publish

  • WebMCP

    A proposed way for a page to expose actions to a browser agent, so an agent that is already on the page can call a typed action instead of clicking around. Chromium-centric as of autumn 2026. It covers acting after arrival. It does not affect whether an assistant names you.

    See: Checkout in chat comes after the shortlist

  • Agentic commerce

    The umbrella term for AI systems acting on a person’s behalf in a shopping context. Two different jobs get sold under the same phrase: getting recommended by a third-party assistant (ChatGPT, Perplexity, Gemini), and making your own site’s search or catalog directly queryable by an agent that already arrived with intent. Under that umbrella sit live catalog lookup (MCP), checkout (ACP/UCP), on-site agent-native search, and a browser agent that uses the page itself. That browser route sits beside live lookup and checkout, and it still does not get you named. Ask which specific piece a vendor means before believing “agentic-commerce-ready” implies anything about being recommended.

    See: Checkout in chat comes after the shortlist. Your catalog can be perfect and still never show up. The assistant said the table was booked

  • ACP (Agentic Commerce Protocol)

    OpenAI and Stripe’s open protocol (beta) for completing a purchase inside a chat — cart, payment, order — once a shopper has already chosen a product. OpenAI has since added product-feed paths for discovery; still not the same as being named on a stranger’s shortlist.

    See: Checkout in chat comes after the shortlist

  • UCP (Universal Commerce Protocol)

    Google’s open standard for a broader shopping interaction — discovery through post-purchase — on AI Mode, Gemini, and partner surfaces. Checkout infrastructure and catalog pipes, not a recommendation mechanism.

    See: Checkout in chat comes after the shortlist

09

Places and travel 3 terms

Restaurants, hotels, walking tours, and trips: different directories, wrong hours, and a bias toward famous places.

In the reading list: Places and travel — start with Why you got the same five landmarks

  • Digital overtourism

    An algorithmically driven concentration of destination visibility before anyone travels — the same iconic places named again and again in AI travel answers, which can preload demand at spots that were already crowded. Distinct from physical overtourism on the ground, but related.

    See: Why you got the same five landmarks

  • Local directory

    The place database a specific chatbot uses for local answers — Foursquare and Yelp for ChatGPT search, Yelp and Tripadvisor for Perplexity, Google Maps for many Gemini answers.

    See: AI place answers come from different directories

  • LocalBusiness

    A Schema.org type for a physical place on your own webpage — name, address, phone, hours — usually as JSON-LD. Prefer a specific subtype (Restaurant, Store) over the generic type. Hours and address have to match your Google Business Profile and other listings; stale markup is the same failure as stale directory data. ChatGPT and Perplexity often lean on licensed directories instead of your site markup.

    See: AI place answers come from different directories

10

Adjacent terms 3 terms

Words you will see elsewhere; not all have a dedicated article yet.

Lookup only — terms you will see elsewhere; not every entry has a dedicated article yet.

  • Hallucination

    When a model states something confidently and fluently that isn’t true — a wrong price, a nonexistent place, a discontinued SKU described as current. Almost every failure mode this site documents is a specific, diagnosable cause of hallucination rather than a mysterious glitch.

    See: Why AI won’t just say it doesn’t know

  • LLM (large language model)

    The model inside an AI chatbot or assistant, trained on very large amounts of text so it can predict and write language. ChatGPT, Claude, Gemini, and similar products are built on LLMs.

  • AI watermarking / content provenance

    Technical methods for marking AI-generated content so it can be identified as synthetic — C2PA metadata, statistical signals like SynthID, and similar. The EU AI Act’s Article 50 transparency rules apply from 2 August 2026 for providers serving EU users; marking is imperfect and often stripped in transit. Relevant background for fake media and deepfakes.

    See: AI can give stupid or dangerous advice

Missing something you'd expect to see here? That probably means an article assumes a term it never defines — tell us.