Resources
Resources
Papers, official docs, and specs cited in the articles. Words we define are on the Glossary. 77 sources in 7 sections.
01
Research & papers 19 sources
Studies the articles rely on: how answers get built, why vague place and product asks repeat the same names, why travel lists repeat, why a model sounds sure, and when a browser agent can open a page and still fail the task.
Research & papers 19 sources
- GEO: Generative Engine Optimization Peer-reviewed definition of GEO and which content tactics moved citation visibility in their tests. arxiv.org
-
RAG and retrieval
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks Original RAG paper (Lewis et al., NeurIPS 2020) — what a model already stores versus what it looks up for the question. Origin of the term. arxiv.org
- What is retrieval-augmented generation (RAG)? — IBM Think Maintained written explainer: training cutoff vs live lookup, retrieve → augment → generate, and why documents get split into chunks. Enterprise-flavored — the pattern is the part we use. ibm.com
- What is Retrieval-Augmented Generation (RAG)? — IBM Technology (video) Optional ~7 min walkthrough of the same three beats if you prefer watching. youtube.com
- Deep-Research Agents Can Be Poisoned via User-Generated Content Cornell Tech paper on WARP — retrieval overlap on UGC pages (STORM, Co-STORM, OmniThink). Exposure vs citation vs mention. arxiv.org
-
robots.txt gatekeeping
- Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web Saarland / WWW 2026 — reputable news sites vs misinformation sites in robots.txt. 60.0% vs 9.1% disallow at least one AI crawler (4,079 sites). arxiv.org
-
Travel and place recommendations
- ChatGPT and the tourist trail: pathway to overtourism or sustainable travel? Mellors (2025) — ChatGPT travel suggestions gravitate to heavily visited destinations unless the asker pushes for alternatives. doi.org
- Destination (Un)Known: Auditing Bias and Fairness in LLM-Based Travel Recommendations Andreev et al. (MDPI *AI*, 2025) — persona-based audit of ChatGPT-4o and DeepSeek-V3 travel recommendations; quantified bias families. Not a universal “all AI” ranking study. doi.org
- Digital overtourism in AI travel recommendations: evidence from a comparative analysis *Current Issues in Tourism* (2026) — coins **digital overtourism**; high nominal variety but low effective diversity across ten AI systems. doi.org
- Assumptions and Undeclared Selection Criteria: The Usefulness of Generative AI as a Travel Recommender System Spennemann (2026) — repeated ChatGPT 5.2 prompts for German Christmas markets narrow to a small canonical set. doi.org
- The Concentrated City: Effects of AI-Generated Travel Advice on the Spatial Distribution of Tourists Barcelona case — ChatGPT heritage picks more concentrated than Instagram geotag spread; supporting point only. doi.org
-
Cold start and prompt context in recommenders
- Revealing Potential Biases in LLM-Based Recommender Systems in the Cold Start Setting Andre, Roy, Dyer, and Wang (Penn / Georgia Tech, 2025). Gemma 3 and Llama 3.2 re-ranking music, movies, and colleges from a provided catalog. A stated task-relevant preference moved lists more than a demographic label. Workshop paper + arXiv preprint. arxiv.org
-
Answer homogeneity across models
- Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond) Jiang et al. — Infinity-Chat, 26K open-ended queries, 70+ models. Intra- and inter-model homogeneity in creative/open-ended answers; English-only. Not a shopping-recommendation study. arxiv.org
-
Hallucination, abstention, and RLHF
- Why Language Models Hallucinate Kalai, Nachum, Vempala, Zhang (Sept 2025) — benchmarks reward guessing over abstention; singleton-rate bound on generative error. Not a ChatGPT ranking guide. arxiv.org
- Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty Zhou et al. — OpenAssistant reward model penalizes hedging language; RLHF step shifts toward confidence. arxiv.org
- Towards Understanding Sycophancy in Language Models Sharma et al. (Anthropic, ICLR 2024) — preference data can favor matching user belief over truth. arxiv.org
- Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models Liang et al. — indifference to truth vs lying; RLHF increases measured "bullshit index" in their tests. arxiv.org
- How RLHF Amplifies Sycophancy Shapira, Benade, Procaccia — formal amplification of slight agreeableness bias through RLHF optimization. arxiv.org
-
Agents using a page
- How AI Agents Use Websites — Merj Lab evidence that an assistant can open a page and still fail the task. The Cancel result is one test and version-specific. The author sells monitoring. The piece is a vendor write-up, not peer reviewed. Accessibility in it is for people first. merj.com
02
Reporting 13 sources
News the articles use: planted forum threads, bad advice, blank product asks and incumbent defaults, a fake brand, and publisher lawsuits.
Reporting 13 sources
-
Planted Reddit and GEO influence
- Companies Are Using Reddit to Manipulate ChatGPT and Google AI Search — 404 Media Primary investigation of r/Biohackers and planted Reddit content aimed at AI answers (Jason Koebler, 3 June 2026). 404media.co
- Chatbots Are the New Influencers Brands Must Woo — The New York Times Erin Griffith (17 Feb 2026) on brands trying to win chatbot mentions. Cited in the Cornell paper’s bibliography. nytimes.com
- How Businesses Are Manipulating ChatGPT Results — The Wall Street Journal Christopher Mims (30 Jan 2026) on GEO/AEO as a paid influence game. Cited in the Cornell paper’s bibliography. wsj.com
-
Dangerous or stupid advice
- Google Is Paying Reddit $60 Million for the glue-on-pizza AI Overview joke — 404 Media Primary trace of the 2024 glue-on-pizza AI Overview back to an 11-year-old Reddit joke. 404media.co
- 3 Hikers Rescued After AI Plan Trip — ABC News Mount Shasta rescue (Aug 2026) — hikers relied on Gemini for packing and route; sheriff on the record. abcnews.com
- AI is having a field day with misinformation sites — Fast Company Interview with Steinacker-Olsztyn on crawling as “do now and ask for forgiveness later.” Journalism around the Saarland paper, not the finding. fastcompany.com
-
Blank product asks and incumbent defaults
- We’re all wearing Uniqlo to the Singularity — Deana Burke (X article) Burke’s fashion probe — seven models, ~70 API answers, no preferences; Uniqlo on 94% of blank white-t-shirt asks. Pooled snapshot; companion to Agent Shelf methodology. x.com
- Agent Shelf (Regular Benchmarks) Ongoing monthly “what should I buy” panel (Claude, ChatGPT, Gemini); methodology explains 10-trial snapshots vs 100-trial cells and why pooled cross-model rates were dropped. regularbenchmarks.com
- Shopify (SHOP) Q2 2026 earnings call transcript — The Motley Fool Harley Finkelstein — 75% of AI-attributed orders outside top 100 categories; agent catalog queries vs keyword search. fool.com
- AI-referred US shoppers browse longer, spend more per visit — Reuters Adobe Analytics (May 2026) — AI-referred retail visits: higher revenue per visit and conversion vs non-AI sources. Stakes context, not a naming guide. reuters.com
-
Fake brand on the shortlist
- The fake brand that ChatGPT fell for — Boys Club (Malware) Deana Burke’s primary newsletter write-up of the Morrowen experiment (5 Aug 2026) — $11.25 domain, July naming window, agentic-commerce stakes. boysclub.beehiiv.com
- A Marketer Built a Fake Deodorant Brand — Inc. Jelinda Montes (8 Aug 2026) on Burke’s Morrowen test — no Reddit/UGC, nearly 9,000 prior “what to buy” queries, Claude and Gemini never named it. inc.com
- How a fake deodorant brand got recommended by AI — Marketplace Kristin Schwab interview with Burke (8 Sept 2026) — public-radio business reporting on the Morrowen experiment and the “AI shelf.” marketplace.org
03
Official platform docs 27 sources
Docs from the companies themselves: crawler names, robots.txt, Cloudflare controls, product feeds, and Search Console.
Official platform docs 27 sources
-
Plans and pricing
- ChatGPT pricing Official Free / Go / Plus / Pro comparison for the consumer chat. Names move; the ordinary-chat ladder (Luna, Sol, Astra) is the part this site uses. chatgpt.com
-
AI crawlers
- Overview of OpenAI Crawlers GPTBot vs OAI-SearchBot vs ChatGPT-User — training, search index, and user-triggered fetch. developers.openai.com
- Google's common crawlers (Google-Extended) Google-Extended token for Gemini training and grounding — separate from Search indexing. developers.google.com
- Perplexity crawlers PerplexityBot (indexing) vs Perplexity-User (live fetch on user questions) — different jobs, different robots.txt rules. docs.perplexity.ai
- Anthropic web crawlers (ClaudeBot, Claude-SearchBot, Claude-User) Official split — training (ClaudeBot), search index (Claude-SearchBot), user-directed fetch (Claude-User) — and how robots.txt opt-out is interpreted. support.anthropic.com
-
robots.txt
- robots.txt — Google Search documentation Plain-language intro to what robots.txt is for crawlers; pairs with the glossary definition. developers.google.com
- Robots Exclusion Protocol (RFC 9309) Normative spec for how a site asks crawlers which pages they may fetch. It does not control access. rfc-editor.org
-
Cloudflare
- Your site, your rules: new AI traffic options for all customers — Cloudflare Official 1 July 2026 post — Search / Agent / Training crawler categories and 15 September 2026 defaults. Primary source, not news coverage. blog.cloudflare.com
- Have it both ways: stay discoverable in search while disallowing AI training — Cloudflare 15 September 2026 launch — Disallow AI Training vs Block, Accountable mixed-use crawlers, and new onboarding defaults. blog.cloudflare.com
- Content Independence Day, one year on — Cloudflare June 2026 bot-mix data — training vs search crawl purposes, mixed-use crawlers. Primary source for machine-traffic composition. blog.cloudflare.com
- Cloudflare Radar Industry bot vs human traffic baselines — context for crawl-to-refer ratios and HTML vs all-HTTP slices. radar.cloudflare.com
- Introducing Pay Per Crawl — Cloudflare Official July 2025 announcement — HTTP 402 / payment headers for AI crawler fetches (closed beta). blog.cloudflare.com
- Trapping misbehaving bots in an AI Labyrinth — Cloudflare March 2025 launch post — decoy pages and invisible links for non-compliant crawlers. blog.cloudflare.com
- AI Labyrinth — Cloudflare Docs Official product documentation for the honeypot feature. developers.cloudflare.com
-
Structured data and feeds
- Intro to structured data (Google Search) Plain starting point for what structured data is for Search. developers.google.com
- Product structured data (Google) Practical Product / Offer JSON-LD fields for commerce pages. developers.google.com
- Rich Results Test (Google) Validate Product/Offer JSON-LD on a live URL before you trust what machines can read. search.google.com
- Generative AI features on Google Search (Google) Google’s line on AI Overviews and AI Mode — no special schema required; ordinary structured data still helps rich results on Search. developers.google.com
- Optimizing content for AI search answers (Microsoft Advertising) Microsoft’s Copilot-oriented note that schema helps search engines and AI systems understand content — not a guarantee of citation or naming. about.ads.microsoft.com
- Local business structured data (Google) Practical LocalBusiness / subtype fields for place pages (hours, address, phone). developers.google.com
- Google Merchant Center — product data specification Feed “dictionary” — core attributes like `id`, `title`, `price`, `gtin`, `mpn`, `condition`. support.google.com
- About product data (Merchant Center) How product data works in Merchant Center (concept, not full account setup). support.google.com
- Google Skillshop (Shopping modules) Free Shopping / Merchant Center learning path (optional certification lane). skillshop.exceedlms.com
-
Search Console
- Web multimodal search performance reporting — Google Search Central Blog Launch post for the Web: multimodal filter — Lens, Circle to Search, image upload to Search, and Chrome “Search this image.” Measures Google web traffic when an image was part of the search, not whether ChatGPT named a shop. developers.google.com
- Performance report (Search results) — Search Console Help Web: text-based is typed queries; Web: multimodal is web results where an image was part of the search. The Image search type is a separate report from multimodal web results. support.google.com
- Performance report: Dimensions and data groupings — Search Console Help The Queries tab is unavailable for multimodal because those searches mostly use images rather than text. Pages, country, and device still group the rows. support.google.com
- Generative AI performance report (Search) — Search Console Help Same Web: multimodal filter for impressions of your links in AI Overviews and AI Mode. That report does not show clicks or query text. support.google.com
04
Specs & conventions 4 sources
Shared formats a page can publish: Schema.org product types, and the proposed llms.txt file.
Specs & conventions 4 sources
- Schema.org — Product Vocabulary for product entities (name, brand, offers, etc.). schema.org
- Schema.org — Offer Price, currency, availability on an offer. schema.org
- Schema.org — Restaurant LocalBusiness subtype for restaurants — hours, address, cuisine-related fields. schema.org
- llms.txt Proposed Markdown map for agents. It is not a crawl-permission file and not a ranking switch. llmstxt.org
05
Agent & commerce protocols 8 sources
Live catalogs and checkout in chat are separate from being named on a shortlist.
Agent & commerce protocols 8 sources
- What is the Model Context Protocol (MCP)? Open standard for connecting AI apps to external data and tools in-session. It is not a public-site ranking switch. modelcontextprotocol.io
- MCP — Architecture overview Host, client, and server roles; tools, resources, and prompts; local and remote transports. modelcontextprotocol.io
- MCP — Tools (specification) How servers expose callable tools and how clients list and invoke them. modelcontextprotocol.io
- MCP — Build a server Official walkthrough for creating a server and connecting a host app (example uses Claude Desktop). modelcontextprotocol.io
- Agentic Commerce Protocol — OpenAI Developers OpenAI’s overview of ACP — catalog ingestion and checkout in agentic commerce flows. developers.openai.com
- Agentic Commerce Protocol — Stripe Stripe’s ACP documentation — checkout, cart, and payment in agent-driven commerce. docs.stripe.com
- Universal Commerce Protocol (UCP) Open standard for agentic purchase flows across platforms — discovery through checkout. ucp.dev
- Google Universal Commerce Protocol (UCP) Guide Google’s guide to adopting UCP on AI surfaces such as Gemini and AI Mode in Search. developers.google.com
06
Learning roadmaps 3 sources
Useful maps of the field. Treat tactics as hypotheses to test, not automatic gospel.
Learning roadmaps 3 sources
- Learning AI Search (Aleyda Solís) Free, maintained AI-search / AEO landscape. Useful map. Treat tactics as hypotheses to test, not gospel. learningaisearch.com
- The AI Search Optimization Checklist (Aleyda Solís) A maintained measurement process: what to log, and that appearing, being recommended, being linked, winning a comparison, and being described accurately are separate outcomes. A consultant’s workflow, not evidence. Her linked studies are not cited here until opened. aleydasolis.com
- WordLift blog (entities) Readable writing on entities vs keywords / knowledge graphs. Skim for concepts; we are not endorsing a vendor stack. wordlift.io
07
This site 3 sources
We dogfood the access layer we write about. That is not a ranking trick.
This site 3 sources
- Fetch check Free check whether AI bots can open your site (robots.txt, llms.txt, sitemap, structured data, thin HTML). Not a recommendation score. This site
- robots.txt Crawler permissions for this site. This site
- llms.txt Our curated map for agents (same idea we explain in the access article). This site