Reference · Glossary

AI search glossary

Definitions for the terms used throughout this site's research — kept tight, each linked to the fuller page where the concept is actually explored.

Ritik Namdev Ritik Namdev ·Published September 2026 ·64 terms ·15 min read
How to use this page

Definitions grouped by what they describe: the disciplines, the surfaces, the bots, the metrics, the mechanisms, and the research vocabulary. Where a term has a full deep-dive elsewhere on this site, the section text links to it. If you only read one section, make it the bots one. That is where the costliest confusion lives.

Terms get added as this site's own research introduces or clarifies them. This list reflects where that research stands today. It is not a fixed, one-time compilation.

The disciplines and their names

Five labels for overlapping ideas, none of them governed by a standards body. The overlap is real and the inconsistency is not your imagination. See the terminology comparison for how each term's scope starts and ends, and which one is actually gaining ground.

AI SEO

The umbrella term for optimizing a site across the whole AI-search landscape. It covers generative answers, AI crawler access, and the technical work behind both. It names a category, not a single tactic. Anyone who says they "do AI SEO" has told you almost nothing actionable yet.

GEO (Generative Engine Optimization)

Writing content so AI systems cite it. Think ChatGPT, Perplexity, Gemini, Claude. A 2023 paper from Princeton, Georgia Tech, the Allen Institute, and IIT Delhi coined the term. It remains the field's only peer-reviewed controlled study, which is a large part of why this term has more weight than its rivals.

AEO (Answer Engine Optimization)

The older, broader term. It covers any answer surface: featured snippets, voice assistants, and generative AI answers alike. It predates GEO by several years, emerging around the rise of Siri, Alexa, and Google's featured snippets. Use it when your scope genuinely includes non-generative answer boxes.

LLMO (Large Language Model Optimization)

A more technical name for GEO. It names the model instead of the answer surface. Fewer people use this term than GEO. It is arguably the more precise label, but precision did not win adoption here.

ASO (Agent Store Optimization)

A proposed term, not yet an established one. It means optimizing a product so an AI shopping agent considers it. Distinct from GEO, which is about answer citation, and from classic SEO, which is about search ranking. No rigorous methodology has been published for it yet.

The engines and surfaces

Coverage routinely treats these as interchangeable. They are not. Conflating AI Overviews with AI Mode, or either with ChatGPT Search, is how a statistic measured on one surface ends up quoted as if it described all of them.

AI Overview

A generated summary that appears above classic organic results on Google for a qualifying query. It synthesizes an answer from multiple retrieved sources and attaches citations. It is not the same as a featured snippet, which extracts one passage verbatim without generating new text.

AI Mode

Google's dedicated, conversational search experience. A separate product from AI Overviews, though both reportedly use query fan-out. Coverage frequently treats the two as interchangeable, which produces conflated statistics.

ChatGPT Search

OpenAI's live web-search capability inside ChatGPT. It retrieves pages at query time rather than answering purely from training data, and attaches citations to search-derived claims. Historically it has drawn in part on a Bing-derived index.

Answer Engine

Any system that returns a direct answer rather than a ranked list of links. The category spans generative chat interfaces, AI Overviews, featured snippets, and voice assistants. Useful when you need a word that covers all of them at once.

Retrieval Layer

The part of an AI system that finds and fetches source material before an answer gets written. Distinct from the model that composes the answer. Most GEO work targets the retrieval layer, not the model.

Crawlers and bots

The most consequential section here. Every major AI company now runs several distinct bots with different jobs, and the decision to allow or block them is per-bot, not one switch. Full detail, including robots.txt syntax and verification methods, is in the AI Bot Registry.

GPTBot

OpenAI's training crawler. It gathers content to improve future models. It has no direct bearing on today's ChatGPT Search citations. Blocking it is a content-policy decision, and it does not remove you from ChatGPT Search.

OAI-SearchBot

OpenAI's live-retrieval bot for ChatGPT Search. This is the crawler that actually produces citations in a ChatGPT Search answer. Blocking it removes you from ChatGPT Search results. This is the single most consequential bot distinction in AI SEO.

ChatGPT-User

Fetches a page only when a person explicitly references its URL in a ChatGPT conversation. User-initiated rather than autonomous. Generally safe to allow, and robots.txt rules may not apply to it in the same way.

ClaudeBot

Anthropic's training crawler. Like GPTBot, it collects content for future model training and does not produce today's citations.

Claude-SearchBot

Anthropic's live-retrieval bot for Claude's web search tool. This is the one that produces Claude citations. Allow it if you want to appear in Claude's answers.

PerplexityBot

Perplexity's crawler for indexing and live answers. Perplexity declares no separate training crawler, saying it does not train foundation models. Note the documented compliance dispute covered in the AI Bot Registry.

Googlebot

The classic web crawler. It also feeds AI Overviews and AI Mode, which is why there is no way to appear in Google Search while excluding yourself from AI Overviews. It is all one index.

Google-Extended

Not a crawler. A control token that lets you opt out of Gemini and Vertex model training while remaining in Google Search. Distinct from blocking Googlebot, which would remove you from search entirely.

Training crawler

Any bot that collects content to build or fine-tune a model. Blocking one affects whether your content shapes a future model. It has no effect on whether that company's product can cite you today.

Retrieval bot

Any bot that fetches a page live, at query time, to answer a specific question and cite the source. Blocking one removes you from that engine's citations immediately. This is the category to allow if visibility is the goal.

Agentic fetcher

A bot that requests a page only when a user explicitly references or pastes its URL. The least consequential of the three bot categories for visibility strategy, and the least likely to need blocking.

Verified bot

A request a network operator has confirmed comes from a declared, legitimate crawler, usually via reverse-DNS plus a forward lookup. User-agent strings alone are trivially spoofed, so verification matters before you trust a log line.

Metrics and measurement

Almost none of these have an agreed industry definition yet, which is exactly the problem the measurement standard was written to address. When you see any of these numbers quoted, ask which query set, which engines, and over what window.

AI Citation

An AI answer that names or links to a specific source. Definitions vary by engine and by measurement tool. Some require a hyperlink, others accept an unlinked brand mention. That inconsistency is a large part of why cross-tool comparison currently fails.

Citation Rate

The share of tracked queries where a domain gets cited at least once. Measured across a fixed query set and time window. Both qualifiers matter: a citation rate without a stated query set and date is unfalsifiable.

AI Share of Voice (SOV)

A domain's citations as a share of all citations for a query set. This lets you compare competitors on the same fixed base. It answers a different question than citation rate, which measures presence rather than relative share.

Citation Half-Life

The number of days until a cited URL's citation rate falls to half its peak. A new metric, defined for the AI Citation Index. Nobody has measured it publicly before, because every existing study reports a snapshot rather than tracking the same citations forward.

Citation Stability

How consistently the same query returns the same cited sources across repeated runs. AI answers are non-deterministic, so a single run mixes real signal with noise. Reporting stability alongside a citation rate tells you how much of the number to trust.

Cross-Platform Concordance

A measure of how much two AI engines agree on what to cite for the same query set. Uses a Jaccard index: the intersection of cited domains divided by the union. A score of 1.0 means perfect agreement; 0 means no overlap at all.

Citation Concentration

How many distinct domains get cited across a query set. Low concentration means citations spread across many sources. High concentration means a few domains dominate. Useful for judging whether a niche is winnable or locked up.

Crawl-to-Referral Ratio

How many pages an AI crawler reads for every visitor it sends back. A 100:1 ratio means 100 pages crawled per single visitor sent. It measures traffic economics, not citation value, so a high ratio is not automatically a reason to block a bot.

Crawl-to-Citation Latency

The number of days between an AI retrieval bot first fetching a new URL and that URL first appearing in a citation. Tracked per engine, since each runs its own crawler on its own schedule.

Zero-Click Search

A search that ends with no click at all, organic or paid. This predates AI Overviews by years, driven originally by featured snippets, knowledge panels, and calculators. AI Overviews sped it up. They did not create it.

Retrieval mechanisms

How the systems actually find and read source material. Understanding these makes most GEO tactics stop feeling arbitrary, because each tactic is aimed at a specific step in this pipeline.

Query Fan-Out

Google's technique for splitting one query into several parallel sub-searches before writing an answer. AI Overviews and AI Mode use it. Documented in US Patent 11663201B2, filed 2018 and granted 2023, which describes eight sub-query type categories.

Sub-query

One of the narrower searches a fan-out generates from a single user question. A page can be pulled into one sub-query result set without appearing in the others, which is why covering adjacent angles of a topic may matter.

Grounding

Tying a generated answer to retrieved source material rather than relying on the model's own training. Citations are the visible output of grounding. An ungrounded answer has no sources to cite.

Extractability

How easily a system can lift a clean, self-contained passage from a page. Not a formal industry term, but the property most GEO tactics actually target. A page can be highly authoritative and still poorly extractable if its answer is buried.

Two-wave indexing

A pattern where an initial crawl indexes raw HTML and a second, delayed pass renders JavaScript-dependent content. Documented for Googlebot. It means JS-rendered content can lag raw-HTML content by a meaningful margin even where rendering happens at all.

CSR / SSR / static generation

Client-side rendering builds page content in the browser after load. Server-side rendering sends fully-formed HTML. Static generation bakes HTML at build time. A bot that does not execute JavaScript sees full content under SSR or static generation, and an empty shell under pure CSR.

Content and quality concepts

Mostly inherited from classic SEO and extended, sometimes without evidence, into AI citation. Note how many carry a Hypothesis grade rather than a confirmed one.

E-E-A-T

Experience, Expertise, Authoritativeness, Trustworthiness. A framework from Google's own rater guidelines, built for human ranking evaluation. People sometimes extend it to AI-citation strategy, though no direct evidence confirms that extension. It began as E-A-T; the Experience dimension was added later.

YMYL (Your Money or Your Life)

Content where a mistake carries real consequences: health, finance, legal, safety. Google's rater guidelines call for extra scrutiny on these topics. Whether AI engines measurably cite differently for YMYL queries is an open, under-studied question.

Content Freshness (AI-citation context)

The idea that updating a page raises its AI citation rate. Widely repeated as GEO advice. No controlled study has actually tested it yet. It may still be worth doing for accuracy and classic-SEO reasons independent of AI citation.

Thin Coverage (gap-topic)

A topic with no dedicated Wikipedia article, or very little Reddit discussion. Used to study what an engine cites when its usual go-to source has nothing relevant. Most B2B and niche commercial topics fall into this category.

Answer-first structure

Writing so the complete answer appears in the opening sentences, before any preamble. Both a readability practice and an extractability one. A buried answer is frequently an uncited answer.

Unlinked brand mention

An occurrence of a brand or entity name in web content without a hyperlink. Reported to correlate more strongly with AI Overview visibility than backlinks do, though the finding remains correlational and confounded.

The agentic and emerging layer

Newer than the rest, and moving fast. Several of these describe technology that shipped within the last year, so definitions here are more likely to shift than those above.

Agentic Browser

A browser, or browser mode, with an AI agent built in. It can navigate, summarize, and act on web pages for a user. ChatGPT Atlas and Perplexity Comet are examples. Different from a chatbot's web search, because it interacts with a live page rather than a search index.

MCP (Model Context Protocol)

A standard that lets an AI application connect to outside tools and structured data. This is different from an agent just reading a rendered web page. A business can expose an MCP server so agents query its data directly.

WebMCP

An early-preview web standard from Google and Microsoft. It lets an agent interact with a site's own functions directly in the browser. This differs from MCP, which connects an AI app to outside tools generally. Shipped in Chrome Canary in early 2026.

Agentic commerce

Purchases where an AI agent, not a human clicking through a storefront, discovers, evaluates, and completes or initiates a transaction. Coverage often blurs agents that merely recommend with agents that actually complete a purchase.

llms.txt

A proposed Markdown file at a site root that gives AI models a curated map of key pages. Google has said it does not use it, and large-scale studies found most files never get requested. Cheap insurance, not a proven lever.

Research and evidence vocabulary

This section exists because most AI-search content skips it. Without this vocabulary, a vendor-reported correlation and a peer-reviewed causal finding look identical on the page. They are not remotely the same thing, and knowing the difference is what makes the rest of this site's grading system readable.

Provenance Grade

A label showing how well a specific number can be traced to its source: Traceable, Partial, or Broken. Used throughout this site instead of treating every statistic as equally solid. Traceable means you can follow it to a named, checkable original.

Evidence Tier

A label showing how solid a claim is: Fact, Evidence, Hypothesis, or Open question. Different from a Provenance Grade, which rates a specific number's sourcing rather than a claim's overall strength.

Pre-registration

Publishing a study's design, hypothesis, and analysis plan before collecting any data. Standard in science, almost unheard of in SEO research. It is what prevents a null result from quietly disappearing.

Null result

A finding that an expected effect does not appear. Just as useful as a positive finding when the study was well-designed, and far less likely to get published. Knowing a tactic does not work saves real effort.

Publication bias

The pattern where positive, exciting findings get published while negative ones quietly disappear. In AI-search research this is close to total, which distorts the apparent evidence for nearly every tactic.

Correlation vs. causation

Correlation means two things move together. Causation means one produces the other. Almost every published AI-search finding is correlational. Treating one as the other is the most common analytical error in this field.

Confound

A third factor that could explain an observed relationship. Sites that add schema also tend to invest in content quality and technical SEO, so an observed schema-citation link may reflect those instead. Randomization is what removes confounds.

Denominator trap

Quoting a within-category percentage as if it described a whole. "ChatGPT drives over half of AI search traffic" is true within the AI-platform category and badly misleading if read as half of all traffic.

Classic SEO terms that still matter

A lot of AI-search content implies the old vocabulary is obsolete. It is not. These terms still describe real things, and several of them gate whether any AI-search work can succeed at all.

Crawlability

Whether a bot can reach and fetch a page at all. The oldest concept in technical SEO and still the gate on everything else. The bot list changed for AI search. The principle did not.

robots.txt

A plain-text file at a site root that declares which crawlers may fetch which paths. Honoured on the honour system: reputable bots comply, some documented crawlers do not. It is a request, not a wall.

Structured data / JSON-LD

Machine-readable labels describing what page content means. Well-evidenced for Google rich results. Its effect on AI citation is contested: live-fetch tests found systems reading only visible HTML, ignoring JSON-LD entirely.

Domain authority

A vendor-calculated estimate of a domain's overall ranking strength, not a Google metric. Correlates with AI citation more weakly than with classic ranking, which is part of why newer sites have a better shot at citation than at top-10 placement.

Featured snippet

A passage extracted verbatim from one source and displayed above Google results. Distinct from an AI Overview, which generates new text synthesized from multiple sources. The two get conflated constantly in coverage.

IndexNow

A protocol for notifying search engines the moment content changes, so pages can be crawled within minutes rather than waiting for a routine re-crawl. Honoured by Bing, Yandex and others. Google tested it and declined to adopt it.

Sitemap

An XML file listing a site's URLs for discovery. Does real, load-bearing work every day, unlike llms.txt, which remains aspirational. Worth getting right before worrying about newer proposed files.

The pairs people mix up most

Five distinctions worth committing to memory, because getting them wrong has real consequences.

Training crawler vs. retrieval bot. The costliest confusion in AI SEO. One shapes future models. One produces today's citations. A blanket block catches both, and only one of them was the intended target.

AI Overviews vs. AI Mode. Two distinct Google products. A figure measured on one does not automatically describe the other, and coverage frequently blurs them.

Evidence Tier vs. Provenance Grade. One rates a claim's overall strength. The other rates whether one specific number can be traced to its origin. A page can carry both, saying different things.

Correlation vs. causation. Nearly every AI-search finding is the first while being written up as though it were the second. This single distinction determines whether a tactic deserves your budget.

Being retrieved vs. being cited. Systems fetch many pages and quote few. Heavy crawler activity with no citations is a quotability problem, not an access problem, and the two need opposite fixes.

How to read a statistic in this field

The vocabulary above is most useful applied. Here is a short procedure for any AI-search number you encounter, including the ones on this site.

Ask who produced it. A peer-reviewed study, a vendor blog post, and an article citing another article are three very different things. Most figures in this field trace back to a small handful of original sources that get re-cited until they look like consensus.

Ask what population it describes. A percentage without a stated sample is unfalsifiable. "58% of clicks lost" means nothing until you know whether that is all searches, only AI-Overview-triggering searches, or a modelled estimate from a consumer survey. Those are three separate real numbers that circulate as one.

Ask whether it is correlational or causal. Almost every figure here is the former, written up in language that implies the latter. This single question changes what the number is worth.

Ask when it was measured. These products change fast enough that a figure from eighteen months ago may describe a system that no longer exists in that form. A number without a date is a number without a shelf life.

Ask whether the method is reproducible. Could someone outside the organization that published it check the work? For most vendor-reported figures the answer is no, because the underlying corpus is the product. That does not make the number wrong. It does mean it cannot be verified, and it should be labelled accordingly.

Any figure that survives all five is unusually strong for this field. Most do not survive two. The full trace of twelve widely-quoted numbers through exactly this process is in the provenance audit.

How to cite this
Namdev, R. (2026). AI search glossary (v1). Retrieved from https://ritiknamdev.com/blog/ai-search-glossary

Published under CC BY 4.0 — reuse freely with attribution.

FAQ

Frequently asked questions

Why do some terms have full pages and others don't?
A term gets its own full page once this site has real research or a real measurement behind it. Terms without one yet still deserve a clear, short definition. This glossary covers both.
How do you decide what counts as a "new" term worth adding?
When this site's own research introduces a concept, or names something the field hasn't named clearly yet, it gets added here. Citation Half-Life and Thin Coverage are both examples of that.
Is Evidence Tier the same thing as Provenance Grade?
No, and mixing them up is a common mistake. Evidence Tier rates how solid a claim is overall. Provenance Grade rates how well one specific number can be traced back to a real source. A claim can be tiered "Hypothesis" while a specific number inside it is graded "Traceable," or the reverse.
Which single distinction matters most in practice?
Training crawler versus retrieval bot. Blocking a training crawler is a content-policy choice with no effect on today's citations. Blocking a retrieval bot removes you from that engine's answers immediately. A copy-pasted "block all AI" rule that catches both is the most common expensive mistake in this field.
Are these definitions industry-standard?
Some are, some are this site's own. GEO, AEO, E-E-A-T and YMYL have external origins. Citation Half-Life, Evidence Tier, Provenance Grade and Thin Coverage were defined here because the field had no precise term for them. Where a definition is ours, the page it links to states the full method.
Why does the glossary include research vocabulary like confound and publication bias?
Because most of what circulates as AI-search fact is correlational, vendor-reported, or both. Without that vocabulary you cannot tell a strong claim from a weak one, and the difference decides where your time goes.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.