"How long until my new page gets cited by ChatGPT or Perplexity?" That's one of the most common practical questions in AI-search SEO. There's no published, disclosed-method answer to it yet. This study tracks new pages across a panel of sites, from first crawl to first citation, and logs the gap in days.
The most practical unanswered question
Almost every research page on this site addresses a conceptual or methodological gap. This one addresses a purely practical one instead. A site owner publishes something new. They want to know roughly how long before it might show up in an AI answer.
Nobody has published a rigorous, disclosed-method answer to that yet. The closest is this site's own single-domain timeline. That's useful, but it describes one site. It isn't a generalizable pattern on its own.
What gets measured
Crawl-to-citation latency is a count of days. Specifically: the days between an AI retrieval bot's first logged fetch of a new URL, and that URL's first appearance in a citation from the matching engine. Each engine gets tracked separately, since each runs its own crawler on its own schedule.
Study design
- 01 Publish new page On a panel site, timestamped
- 02 Log first crawl Per bot, from server logs
- 03 Query the Index Daily, against the fixed query set
- 04 Log first citation Timestamped, per engine
- 05 Compute latency Days from first crawl to first citation
A worked example of one tracked page
Here's the pipeline made concrete, with invented, hypothetical timestamps. A panel site publishes a new article on Day 0. Server logs show OAI-SearchBot's first fetch of that URL on Day 3. Daily monitoring against the Index's query set first detects a ChatGPT Search citation on Day 11.
The crawl-to-citation latency for that page, on that engine, gets recorded as 8 days: Day 11 minus Day 3. Two separate numbers get logged, not just the total. The 3-day crawl delay is one stage. The 8-day citation delay is another. They likely respond to different factors, so keeping them apart matters.
Pre-registered hypotheses
| # | Hypothesis | Prediction |
|---|---|---|
| LT1 | Median crawl-to-citation latency is under 14 days for at least one major engine | Supported |
| LT2 | Latency correlates negatively with existing domain authority — established sites see faster citation | Supported |
| LT3 | Latency varies significantly by engine, with Bing-derived surfaces faster than others | Supported |
'How long until my new page gets cited by ChatGPT?' is one of the most common questions in AI-search SEO. Nobody has published a disclosed-method answer to it. This study is designed to give one.
Share on XWhat might actually speed things up
Nobody has isolated the exact cause yet, but a few candidates keep coming up in general crawler research. A fresh XML sitemap ping tells a bot a new URL exists right away, instead of waiting to be discovered by a routine re-crawl. Internal links from already-crawled, high-traffic pages give a bot an easy path to find the new one sooner.
A site's existing crawl budget matters too. A domain a bot already visits daily gets checked again sooner than one it visits every few weeks. None of these are confirmed causes of faster citation yet, only of faster crawling. This study is designed to test whether crawl speed and citation speed actually move together, or whether they're more independent than they look.
What our own site already suggests
Zero to Cited is this site's own evidence-backed timeline for a brand-new domain. It documented first citations appearing in month 2–3, and broader multi-engine presence by month 4–6.
That's a single data point, not a generalizable finding. But it's a useful prior for what this larger panel study might confirm, or overturn.
How this compares to classic SEO indexing speed
Classic SEO has a comparable, well-studied concept: the time between a new page's publication and its first appearance in Google's index. That has historically ranged from hours, for a well-established, frequently crawled site, to weeks, for a new or low-authority domain.
Suppose crawl-to-citation latency for AI engines follows a broadly similar authority-dependent pattern. That's the direction hypothesis LT2 predicts. It would suggest AI citation discovery inherits much of the same underlying economics as classic search indexing, rather than running on an entirely different timeline. Whether that holds, or AI citation latency behaves meaningfully differently, is exactly what this study is designed to reveal.
How to use this once results exist
Once published, the main practical use is expectation-setting. A new site owner could compare their own crawl-to-citation experience against a published, panel-derived benchmark. That beats guessing, or relying on a single anecdote.
A site whose latency runs far longer than the panel median gets a concrete signal from that comparison. Something in their technical setup or authority profile may be worth investigating. That's a much better starting point than a vague feeling that AI citation "isn't happening yet."
What to do while you wait
Results aren't published yet, but three low-risk steps are reasonable in the meantime. Submit a fresh sitemap ping whenever you publish a new page, so bots learn about it fast. Link to the new page from an already-crawled page with real traffic, so a bot has an easy path to find it.
And check your server logs for the retrieval bots covered in the AI Bot Registry, to confirm a crawl actually happened before assuming citation is the bottleneck. A page that was never crawled can't be cited, no matter how good it is.
Limitations
- Panel sites are volunteers, not a random sample of the web — likely biased toward site owners already engaged enough with AI-search topics to join a research panel.
- "First citation" depends on query coverage — a page could be citable for a query outside the tracked set and this study wouldn't detect it.
Namdev, R. (2026). Crawl-to-citation latency (v1). Retrieved from https://ritiknamdev.com/blog/crawl-to-citation-latency-study Published under CC BY 4.0 — reuse freely with attribution.
Complements Zero to Cited, this site's own single-domain timeline, with a multi-site, statistically generalizable version.