Original research · Pre-registered · Log study

Crawl-to-citation latency

How many days pass between an AI crawler first fetching a new page and that page first appearing in a citation? The most practical question a new site owner asks, and nobody has published an answer.

Ritik Namdev Ritik Namdev ·Published September 2026 ·v0 — panel recruitment ·14 min read ·Last verified September 2026
The short version

"How long until my new page gets cited by ChatGPT or Perplexity?" is one of the most common practical questions in AI search, and there is no published, disclosed-method answer to it. This study tracks new pages across a panel of sites, from first crawl to first citation, and logs the gap in days — plus the share of pages never cited at all. That second number may turn out to be the more useful one. Panel recruitment is open; no interval has been recorded yet.

Research status Recruiting participants

No data has been collected yet. Nothing on this page is a result. Panel sites are being recruited. No publish-to-citation interval has been recorded under the protocol.

What is known
  • Crawl events are directly observable in server logs, and citation appearance is directly checkable.
  • This site's own zero-to-cited log records the sequence for one site, which is where the design comes from.
What is not yet known
  • Typical elapsed time from publication to first citation, across a panel rather than one site.
  • Whether latency differs by engine.
  • Whether it differs by site authority or topic.

Research question: how long from crawl to citation?

The question: how many days pass between an AI retrieval bot's first logged fetch of a new URL and that URL's first appearance in a citation from the matching engine? Each engine is tracked separately, because each runs its own crawler on its own schedule. OpenAI, Perplexity, Anthropic and Google each document their own, and the bot registry keeps them in one table.

Why it is unanswered: answering it needs two things almost nobody has together — durable server logs across many domains, and a fixed query set checked daily against every engine. Vendors have the query side and not the logs; site owners have the logs and not the panel. So the latency figure that circulates is nobody's measurement — a pattern traced in the provenance audit.

What is known today, and from where?

Zero to Cited is this site's own log-backed timeline for a brand-new domain: first citations in month 2–3, broader multi-engine presence by month 4–6. One site, one niche, one technical setup — a prior, not a benchmark, and not a result this study may repeat.

Classic SEO offers the nearest analogue. Time from publication to first appearance in Google's index has historically ranged from hours on an established, frequently crawled site to weeks on a new or low-authority domain, with Ahrefs finding considerably longer again to actually rank. Whether AI citation inherits that same authority-dependent economics is precisely what LT2 predicts and this study tests.

Several mechanisms plausibly move the crawl half of the gap. A fresh XML sitemap tells a bot a URL exists rather than waiting for re-crawl discovery, and Bing has said explicitly that sitemaps still matter for AI-powered search. IndexNow is the push-based version — setup takes minutes, Microsoft reports it drives faster discovery, adoption keeps expanding, and its submission-based ancestor predates the AI wave. That lineage matters because one analysis found 87% of SearchGPT citations matching Bing's top results. That dependency is why the Bing and Copilot side of this is worth treating as a distinct route in rather than a footnote to Google.

Processing cost is the other half. Existing crawl budget sets how often a bot returns, and JavaScript costs roughly nine times the crawl time of HTML, with rendering adding its own delay. That is why the rendering study sits upstream of this one. Hypothesis None of these is an established cause of faster citation — only of faster crawling. Whether the two move together is the open question.

Method: sample, control and variables

Tracking pipeline, per new page
  1. 01 Publish new page On a panel site, timestamped
  2. 02 Log first crawl Per bot, from server logs
  3. 03 Query the Index Daily, against the fixed query set
  4. 04 Log first citation Timestamped, per engine
  5. 05 Compute latency Days from first crawl to first citation

Sample. A recruited panel of sites that publish new pages regularly and can supply durable server logs, deliberately spanning a range of domain authority so LT2 is testable rather than assumed. Each new page is registered with its exact publication timestamp before it goes live.

Control. Pages published across the same window by other panel sites act as the comparison group for any single site's number. An engine-side change during the window then shows up as a shift affecting everything at once, not as one site's improvement.

Variables. Two intervals are logged per page per engine, never one blended figure: publish-to-crawl, from the timestamp to the first verified 200-status bot fetch; and crawl-to-citation, from that fetch to the first observed citation. Recorded alongside: engine, bot user-agent, domain authority band, topic category, and whether the page was pushed via sitemap or IndexNow.

Planned editions, fixed in advance
  1. v0Now

    Design, pre-registration and panel recruitment

    This page. Hypotheses, stage definitions and the commitment to report the never-cited share, all fixed before collection.

  2. v2Second cohort

    Re-run with a wider authority spread

    Tests LT2 properly by deliberately sampling across established and new domains.

  3. v3Later

    Updated pages rather than new ones

    A different question with a likely different answer, registered rather than assumed.

Publishing the design before the data is the discipline used on the robots.txt census. Everything lands in the studies index with its raw rows attached.

Hypothesis: what we predict before running it

Predictions registered before collection. No data has been collected — these are expectations, not results.
#HypothesisPredicted outcome
LT1Median crawl-to-citation latency is under 14 days for at least one major engineExpect to hold
LT2Latency correlates negatively with existing domain authority — established sites see faster citationExpect to hold
LT3Latency varies significantly by engine, with Bing-derived surfaces faster than othersExpect to hold

'How long until my new page gets cited by ChatGPT?' is one of the most common questions in AI-search SEO. Nobody has published a disclosed-method answer to it. This study is designed to give one.

Share on X

Five stages hidden inside one latency number

"Crawl to citation" sounds like one gap. It is a chain, and each link has its own delay and failure mode.

StageWhat has to happenObservable from logs?
DiscoveryThe bot learns the URL exists, from a sitemap, a link or a feed.No — inferred
FetchThe bot requests the page and gets a usable response.Yes
ProcessingThe content is parsed, chunked and stored in whatever index the engine uses.No
Retrieval eligibilityThe page becomes a candidate for a matching query.No
SelectionThe page is actually chosen and shown as a citation.Yes — by querying

Only the second and fifth stages are directly observable; everything between is inferred from the gap. That is an uncomfortable fact about this design, and it belongs up front rather than buried. It is also why two numbers are logged rather than one.

Publish-to-crawlCrawl-to-citation
What it measuresDiscovery and crawl budgetEverything downstream of the fetch
Observable directly?Yes, from logsOnly at both ends
Within your control?Largely — sitemaps, links, server healthBarely — retrieval and selection are the engine's
What fixes a slow numberPush-based submission, internal links, faster responsesA better page, or a less crowded question
What a bad number meansA technical problem you can findA competitiveness problem, or simply a wait

Analysis: the pages that are never cited at all

This is the methodological issue most likely to break a naive version of the study. Track a hundred pages for ninety days; forty get cited, sixty do not. Compute the median from the forty and you have measured the pages that succeeded while quietly deleting the ones that did not. The number will look reassuring and be wrong.

Statisticians call this censoring, and it has established handling: report the share of pages still uncited at each time point rather than only the average among those that were. The output stops being "the median is N days" and becomes "by day thirty, this proportion had been cited at least once" — harder to quote, considerably more truthful.

Hypothesis We expect the never-cited share to be large, possibly the majority. If so, most of the study's practical value lies in that figure rather than the latency. The question then moves back to what makes a page citable at all — the ground covered by the tactic evidence scoreboard and the schema study.

Control: confounds in a latency measurement

A latency number can be moved by things that have nothing to do with an engine's speed. A panel design has to handle each.

Confounds the panel design has to screen for. Each one can inflate a measured latency without any engine being slow. Listed before collection so the screens are fixed in advance, not chosen after seeing the numbers.
ConfoundHow it inflates measured latencyScreen
Query set coverageA page is only detected as cited if a tracked query would surface it, so poor coverage looks exactly like slow citationFixed query set registered per page before publication
Detection frequencyDaily checking rounds latency to whole days and misses citations that appeared and vanished between checksReport the check interval alongside every interval; treat citation as impermanent
Topic competitivenessA page in a crowded topic competes against established sources, often Wikipedia; a page in an empty niche does notRecord topic category; never blend categories into one median
Page quality variationA slow page may simply be a weak page, and no timing measurement can separate thoseReport the never-cited share separately from the latency distribution
Engine-side changeAn index refresh or retrieval change mid-window shifts latency for everything at onceSame-window pages from other panel sites as the comparison group

Engine-side change. An index refresh or retrieval change mid-window shifts latency for everything at once. So does whatever freshness weighting an engine applies, tested directly in the freshness study. Newer fetch modes behave differently again, as 30-day log studies of agentic crawling show.

Why server logs corrupt the start time

The whole measurement rests on knowing when a bot first fetched the page. Server logs make that look easy. Six specific ways it goes wrong, each of which the panel protocol has to screen for before a log line is allowed to become a start time:

User-agent spoofingAny client can claim to be any bot. Verify against published IP ranges before a log line becomes a start time.
Multiple bots per operatorTreating a training fetch as a retrieval fetch corrupts the measurement at its origin.
CDN and cache layersIf the edge serves the bot, the origin never logs the request. The page was fetched and you cannot see it.
Log retention windowsA study window longer than your host keeps logs loses its own start times.
Non-200 status codesA redirect, an error or a rate limit is not a successful fetch. Counting it shortens latency artificially.
Clock and timezone driftComparing a server log against a monitoring system introduces offsets that matter at day-level resolution.

Network-scale context helps sanity-check a panel site's own volumes: Cloudflare's crawler breakdown and crawl-to-click analysis, with independent restatements of the same ratios.

Why latency must be reported per engine

Latency is a per-engine property, and the architectures differ at exactly the point that sets the delay. Some surfaces answer from a live fetch at query time; others from a pre-built index refreshed on its own schedule. Those produce fundamentally different distributions, and an average across them means nothing. Some engines inherit discovery from an established web index, so a page already in that index may become citable almost immediately while an unknown page waits — the asymmetry LT3 exists to test.

Caching adds a layer: an engine may hold a copy from before your update, so what it cites is not what you published. Latency for a new page and for an updated page are different questions, and this study measures the first. The practical rule follows: never quote one crawl-to-citation figure for "AI search". Quote it per engine, with the date. Even inside Google the two surfaces disagree with each other, and the concordance study measures the spread.

Interpretation: how to read the number, and how not to

The intended use is expectation-setting. A site owner compares their own crawl-to-citation experience against a panel-derived benchmark. A latency far above the panel median is a concrete signal that something in the technical setup or authority profile is worth investigating — a better starting point than a vague feeling that AI citation "isn't happening yet". Four readings would be wrong.

"The median is N days, so my page should be cited by then." A median describes a population; half of it took longer. One page is a draw from a wide distribution, not a scheduled event.

"My page was crawled, so citation is coming." Crawling is necessary, not sufficient. Most crawled pages are never cited for any given query.

"Latency went down, so our optimisation worked." Or the engine changed, or the topics did. Without pages published over the same window as a comparison, those cannot be separated.

"Nothing in three weeks, so AI search doesn't work for us." Three weeks may be entirely normal. That sentence gets said so often precisely because no benchmark exists to check it against, and it is how good content programmes get cancelled early.

Null results we would publish

All three hypotheses can fail, and each failure publishes as prominently as a confirmation. Median latency exceeding fourteen days on every engine measured contradicts LT1. No relationship between domain authority and latency contradicts LT2. No meaningful difference between engines contradicts LT3. So does the outcome that would redirect the whole thread — a never-cited share so high that latency is not the interesting variable.

An infrastructure failure counts too: insufficient log quality across the panel to establish reliable start times. Saying "we could not measure this cleanly" is a real contribution in a field publishing numbers without saying how they got them. Each lands in the null results registry, under the standards this site publishes to.

Reproduction: measuring your own latency

A single-site version will not generalise, but it will tell you whether your own pipeline works at all, which is usually the real question. The citation playbook covers the page itself; the free tools here cover the mechanical parts.

A single-site latency measurement you can run yourself
  1. 1 Record exact publication timestamps Somewhere structured, not in your memory.
  2. 2 Set up durable, queryable log access If your host rotates logs weekly, ship them somewhere that keeps them longer than your study window.
  3. 3 Filter for retrieval bots and verify them Check status codes, and confirm the bot is what it claims where IP ranges are published. Record the first successful fetch per bot per URL.
  4. 4 Fix a small query set before you publish Three to five natural questions the page should answer. Written in advance, so you cannot fit them to the outcome.
  5. 5 Check on a sustainable schedule and log the negatives The absence of a citation on a given day is data, and dropping it is how the censoring problem starts.
  6. 6 Report both the distribution and the never-cited share Reporting only the first is the mistake this whole page warns against.

Who this applies to, and who it does not

A latency benchmark is most useful to publishers who ship new pages regularly and can instrument them. Frequency gives you a sample; instrumentation gives you a start time. Without both, the measurement is guesswork dressed as a metric.

It applies differently to a brand-new domain. With no crawl history, the publish-to-crawl stage may dominate everything else, so the crawl-to-citation figure is not the constraint that site is actually hitting — the discovery mechanisms above matter far more first. It applies weakly to sites whose value sits in a handful of stable pages: if you publish four pages a year, watching those four individually beats any distribution.

And it does not apply at all to content that will never be cited for reasons unrelated to timing. Gated material, thin pages and duplicated text have a citability problem that no amount of waiting solves, and reading a slow latency number as an engine problem in those cases points remediation at the wrong thing.

Raw data we will publish

Three files publish with the results. The per-page log extract carries verified first-fetch timestamps by bot. The daily citation-check log includes every negative check. The panel manifest describes each site's authority band, topic category and submission method, without naming sites that asked not to be named. Summary curves are never published without the rows behind them.

Limitations

  • Sample: volunteers, not a random sample of the web. A recruited panel skews toward site owners already engaged enough with AI search to join a research panel, and toward those with the technical maturity to retain logs — both plausibly correlated with faster citation.
  • Geography and language are uncontrolled. Panel sites will not be evenly distributed across regions or languages, and crawl schedules and index refreshes need not be uniform across them. Any result describes the panel's geography, not the web's.
  • Engine coverage is limited to surfaces we can query daily. Engines without a checkable citation list, and newer agentic fetch modes, are outside the design entirely.
  • Query selection bounds the outcome. "First citation" means first citation for a tracked query. A page citable for something outside the fixed set registers as never cited, which inflates the never-cited share by an unknown amount.
  • Measurement resolution is one day. Daily checking cannot resolve sub-day latency, misses citations that appear and disappear between checks, and inherits any clock drift between log and monitor.
  • Confounders survive the design. Page quality, topic competitiveness and publication timing are recorded but not randomised, so no causal claim about what shortens latency can come out of this study.
  • Reproducibility depends on log quality we do not control. CDN edge caching can hide fetches from origin logs entirely, so some start times will be wrong in a direction we cannot detect.
  • Generalisation is time-bounded. A latency figure describes the engines as they behaved during one collection window, and re-running is the only way to know whether it still holds.
Get the results, or join the panel

The curves and the raw log files go out to the newsletter when the first cohort's window closes. If you publish regularly and keep durable server logs, the panel is still being recruited — there is no enrolment form, so the about page explains how to get in touch. The study sits inside the AI Citation Index programme.

How to cite this
Namdev, R. (2026). Crawl-to-citation latency (v1). Retrieved from https://ritiknamdev.com/blog/crawl-to-citation-latency-study

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

Part of the AI Citation Index programme, which supplies the query set. Complements Zero to Cited, this site's own single-domain timeline, with a multi-site version; see also the citation half-life study, which asks how long a citation lasts once won.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

Google Search Central — Build and submit a sitemapdevelopers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap Google Search Central — Managing crawl budget for large sitesdevelopers.google.com/search/docs/crawling-indexing/large-site-managing-crawl-budget Google Search Central — Overview of Google crawlers and fetchersdevelopers.google.com/search/docs/crawling-indexing/overview-google-crawlers Bing Webmaster Blog — Keeping content discoverable with sitemaps in AI-powered searchblogs.bing.com/webmaster/July-2025/Keeping-Content-Discoverable-with-Sitemaps-in-AI-Powered-Search Bing — IndexNow getting startedwww.bing.com/indexnow/getstarted Bing Webmaster Blog — IndexNow drives smarter and faster content discoveryblogs.bing.com/webmaster/May-2025/IndexNow-Drives-Smarter-and-Faster-Content-Discovery Bing Webmaster Blog — IndexNow adoption across industriesblogs.bing.com/webmaster/December-2024/Look-How-Far-We-ve-Come-IIndexNow-Expands-Adoption-Across-Industries Bing Webmaster Blog — Submitting up to 10,000 URLs per dayblogs.bing.com/webmaster/january-2019/bingbot-Series-Get-your-content-indexed-fast-by-now-submitting-up-to-10,000-URLs-per-day-to-Bing Wikipedia — IndexNowen.wikipedia.org/wiki/IndexNow Ahrefs — How long does it take to rank in Googleahrefs.com/blog/how-long-does-it-take-to-rank-in-google-and-how-old-are-top-ranking-pages Onely — Google needs 9x more time to crawl JS than HTMLwww.onely.com/blog/google-needs-9x-more-time-to-crawl-js-than-html Onely — Google rendering delaywww.onely.com/blog/googles-rendering-delay-5-seconds Cloudflare — The crawl-to-click gap for AI botsblog.cloudflare.com/crawlers-click-ai-bots-training Cloudflare — From Googlebot to GPTBot: who is crawling your siteblog.cloudflare.com/from-googlebot-to-gptbot-whos-crawling-your-site-in-2025 Digital Applied — 30-day agentic crawler behaviour log studywww.digitalapplied.com/blog/agentic-crawler-behavior-30-day-site-log-study SEOmator — Crawl-to-refer ratios for AI crawlers and LLM botsseomator.com/blog/crawl-to-refer-ratio-ai-crawlers-llm-bots Momentic — AI search crawlers and bots referencemomenticmarketing.com/blog/ai-search-crawlers-bots OpenAI — Overview of OpenAI crawlersdevelopers.openai.com/api/docs/bots Perplexity — Official crawler documentationdocs.perplexity.ai/docs/resources/perplexity-crawlers Anthropic — Does Anthropic crawl the web, and how to block itsupport.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler Seer Interactive — 87% of SearchGPT citations match Bing top resultswww.seerinteractive.com/insights/87-percent-of-searchgpt-citations-match-bings-top-results Salespeak — Content freshness and AI searchsalespeak.ai/aeo-news/content-freshness-ai-search
FAQ

Frequently asked questions

How long until a new page gets cited by ChatGPT or Perplexity?
Nobody has published a disclosed-method answer, which is why this study exists. This site's own single-domain log records first citations appearing in month 2-3 and broader multi-engine presence by month 4-6, but that is one site, not a benchmark. No interval has been recorded under this protocol yet.
Why measure this across a panel instead of just your own site?
A single site's latency could just reflect that site's authority, technical setup or niche. A panel across many site owners, with varying characteristics, lets the finding generalize rather than describing one domain's quirks.
What happens to pages that are never cited at all? Do they get dropped?
They must not be, and this is the biggest single threat to the study. A latency figure computed only from pages that eventually got cited describes the lucky subset, not the population. We will report the share still uncited at the end of the window alongside any median, because that share is often the more useful number.
Does a faster crawl actually mean a faster citation?
That is exactly what the study is built to find out, and we are careful not to assume it. Crawling and citing are separate stages with separate bottlenecks. A page can be fetched within hours and still take weeks to appear in an answer, or never appear at all.
How is this different from the AI Bot Registry's crawl-to-referral ratio?
That metric measures how many pages a bot crawls per visitor it sends — a traffic-economics measure. This one measures time from crawl to citation for a specific new page: a speed-of-discovery measure. Related, different practical question.
Can I speed this up by publishing more often?
Publishing frequency plausibly affects crawl rate, which is one stage of several. Whether it affects citation is unestablished. The advice has an obvious failure mode: publishing more, worse pages to court a crawler is a reliable way to make a site less citable overall.
How would I know if my own latency is unusually bad?
Right now, you would not, and that is the gap this study exists to fill. Without a published benchmark, a site owner cannot tell a normal wait from a broken pipeline — which is why so much AI-search advice gets bought during weeks that were always going to be quiet.
Will results differ for a page added to an established site versus a brand-new domain?
That is hypothesis LT2, and the panel is designed to include both. Comparing the two is a planned secondary analysis once initial data exists; it is a prediction, not a finding.
Can I join the panel?
Panel sites are being recruited now. There is no enrolment form — the about page explains how to get in touch. What is needed is durable server-log access covering the study window, and a regular publishing cadence so there is a sample to track.
Why publish the design before the study runs?
Because a design published afterwards can be quietly reshaped to fit whatever the data said. Fixing the stage definitions, the hypotheses and the commitment to report the never-cited share in advance is the cheapest protection against moving the goalposts.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.