"How long until my new page gets cited by ChatGPT or Perplexity?" is one of the most common practical questions in AI search, and there is no published, disclosed-method answer to it. This study tracks new pages across a panel of sites, from first crawl to first citation, and logs the gap in days — plus the share of pages never cited at all. That second number may turn out to be the more useful one. Panel recruitment is open; no interval has been recorded yet.
No data has been collected yet. Nothing on this page is a result. Panel sites are being recruited. No publish-to-citation interval has been recorded under the protocol.
- Crawl events are directly observable in server logs, and citation appearance is directly checkable.
- This site's own zero-to-cited log records the sequence for one site, which is where the design comes from.
- Typical elapsed time from publication to first citation, across a panel rather than one site.
- Whether latency differs by engine.
- Whether it differs by site authority or topic.
Research question: how long from crawl to citation?
The question: how many days pass between an AI retrieval bot's first logged fetch of a new URL and that URL's first appearance in a citation from the matching engine? Each engine is tracked separately, because each runs its own crawler on its own schedule. OpenAI, Perplexity, Anthropic and Google each document their own, and the bot registry keeps them in one table.
Why it is unanswered: answering it needs two things almost nobody has together — durable server logs across many domains, and a fixed query set checked daily against every engine. Vendors have the query side and not the logs; site owners have the logs and not the panel. So the latency figure that circulates is nobody's measurement — a pattern traced in the provenance audit.
What is known today, and from where?
Zero to Cited is this site's own log-backed timeline for a brand-new domain: first citations in month 2–3, broader multi-engine presence by month 4–6. One site, one niche, one technical setup — a prior, not a benchmark, and not a result this study may repeat.
Classic SEO offers the nearest analogue. Time from publication to first appearance in Google's index has historically ranged from hours on an established, frequently crawled site to weeks on a new or low-authority domain, with Ahrefs finding considerably longer again to actually rank. Whether AI citation inherits that same authority-dependent economics is precisely what LT2 predicts and this study tests.
Several mechanisms plausibly move the crawl half of the gap. A fresh XML sitemap tells a bot a URL exists rather than waiting for re-crawl discovery, and Bing has said explicitly that sitemaps still matter for AI-powered search. IndexNow is the push-based version — setup takes minutes, Microsoft reports it drives faster discovery, adoption keeps expanding, and its submission-based ancestor predates the AI wave. That lineage matters because one analysis found 87% of SearchGPT citations matching Bing's top results. That dependency is why the Bing and Copilot side of this is worth treating as a distinct route in rather than a footnote to Google.
Processing cost is the other half. Existing crawl budget sets how often a bot returns, and JavaScript costs roughly nine times the crawl time of HTML, with rendering adding its own delay. That is why the rendering study sits upstream of this one. Hypothesis None of these is an established cause of faster citation — only of faster crawling. Whether the two move together is the open question.
Method: sample, control and variables
- 01 Publish new page On a panel site, timestamped
- 02 Log first crawl Per bot, from server logs
- 03 Query the Index Daily, against the fixed query set
- 04 Log first citation Timestamped, per engine
- 05 Compute latency Days from first crawl to first citation
Sample. A recruited panel of sites that publish new pages regularly and can supply durable server logs, deliberately spanning a range of domain authority so LT2 is testable rather than assumed. Each new page is registered with its exact publication timestamp before it goes live.
Control. Pages published across the same window by other panel sites act as the comparison group for any single site's number. An engine-side change during the window then shows up as a shift affecting everything at once, not as one site's improvement.
Variables. Two intervals are logged per page per engine, never one blended figure: publish-to-crawl, from the timestamp to the first verified 200-status bot fetch; and crawl-to-citation, from that fetch to the first observed citation. Recorded alongside: engine, bot user-agent, domain authority band, topic category, and whether the page was pushed via sitemap or IndexNow.
- v0Now
Design, pre-registration and panel recruitment
This page. Hypotheses, stage definitions and the commitment to report the never-cited share, all fixed before collection.
- v1First cohort
New pages tracked from publication to first citation
Two numbers per page per engine: publish-to-crawl, then crawl-to-citation. Raw logs published with the summary.
- v2Second cohort
Re-run with a wider authority spread
Tests LT2 properly by deliberately sampling across established and new domains.
- v3Later
Updated pages rather than new ones
A different question with a likely different answer, registered rather than assumed.
Publishing the design before the data is the discipline used on the robots.txt census. Everything lands in the studies index with its raw rows attached.
Hypothesis: what we predict before running it
| # | Hypothesis | Predicted outcome |
|---|---|---|
| LT1 | Median crawl-to-citation latency is under 14 days for at least one major engine | Expect to hold |
| LT2 | Latency correlates negatively with existing domain authority — established sites see faster citation | Expect to hold |
| LT3 | Latency varies significantly by engine, with Bing-derived surfaces faster than others | Expect to hold |
'How long until my new page gets cited by ChatGPT?' is one of the most common questions in AI-search SEO. Nobody has published a disclosed-method answer to it. This study is designed to give one.
Share on XFive stages hidden inside one latency number
"Crawl to citation" sounds like one gap. It is a chain, and each link has its own delay and failure mode.
| Stage | What has to happen | Observable from logs? |
|---|---|---|
| Discovery | The bot learns the URL exists, from a sitemap, a link or a feed. | No — inferred |
| Fetch | The bot requests the page and gets a usable response. | Yes |
| Processing | The content is parsed, chunked and stored in whatever index the engine uses. | No |
| Retrieval eligibility | The page becomes a candidate for a matching query. | No |
| Selection | The page is actually chosen and shown as a citation. | Yes — by querying |
Only the second and fifth stages are directly observable; everything between is inferred from the gap. That is an uncomfortable fact about this design, and it belongs up front rather than buried. It is also why two numbers are logged rather than one.
| Publish-to-crawl | Crawl-to-citation | |
|---|---|---|
| What it measures | Discovery and crawl budget | Everything downstream of the fetch |
| Observable directly? | Yes, from logs | Only at both ends |
| Within your control? | Largely — sitemaps, links, server health | Barely — retrieval and selection are the engine's |
| What fixes a slow number | Push-based submission, internal links, faster responses | A better page, or a less crowded question |
| What a bad number means | A technical problem you can find | A competitiveness problem, or simply a wait |
Analysis: the pages that are never cited at all
This is the methodological issue most likely to break a naive version of the study. Track a hundred pages for ninety days; forty get cited, sixty do not. Compute the median from the forty and you have measured the pages that succeeded while quietly deleting the ones that did not. The number will look reassuring and be wrong.
Statisticians call this censoring, and it has established handling: report the share of pages still uncited at each time point rather than only the average among those that were. The output stops being "the median is N days" and becomes "by day thirty, this proportion had been cited at least once" — harder to quote, considerably more truthful.
Hypothesis We expect the never-cited share to be large, possibly the majority. If so, most of the study's practical value lies in that figure rather than the latency. The question then moves back to what makes a page citable at all — the ground covered by the tactic evidence scoreboard and the schema study.
Control: confounds in a latency measurement
A latency number can be moved by things that have nothing to do with an engine's speed. A panel design has to handle each.
| Confound | How it inflates measured latency | Screen |
|---|---|---|
| Query set coverage | A page is only detected as cited if a tracked query would surface it, so poor coverage looks exactly like slow citation | Fixed query set registered per page before publication |
| Detection frequency | Daily checking rounds latency to whole days and misses citations that appeared and vanished between checks | Report the check interval alongside every interval; treat citation as impermanent |
| Topic competitiveness | A page in a crowded topic competes against established sources, often Wikipedia; a page in an empty niche does not | Record topic category; never blend categories into one median |
| Page quality variation | A slow page may simply be a weak page, and no timing measurement can separate those | Report the never-cited share separately from the latency distribution |
| Engine-side change | An index refresh or retrieval change mid-window shifts latency for everything at once | Same-window pages from other panel sites as the comparison group |
Engine-side change. An index refresh or retrieval change mid-window shifts latency for everything at once. So does whatever freshness weighting an engine applies, tested directly in the freshness study. Newer fetch modes behave differently again, as 30-day log studies of agentic crawling show.
Why server logs corrupt the start time
The whole measurement rests on knowing when a bot first fetched the page. Server logs make that look easy. Six specific ways it goes wrong, each of which the panel protocol has to screen for before a log line is allowed to become a start time:
Network-scale context helps sanity-check a panel site's own volumes: Cloudflare's crawler breakdown and crawl-to-click analysis, with independent restatements of the same ratios.
Why latency must be reported per engine
Latency is a per-engine property, and the architectures differ at exactly the point that sets the delay. Some surfaces answer from a live fetch at query time; others from a pre-built index refreshed on its own schedule. Those produce fundamentally different distributions, and an average across them means nothing. Some engines inherit discovery from an established web index, so a page already in that index may become citable almost immediately while an unknown page waits — the asymmetry LT3 exists to test.
Caching adds a layer: an engine may hold a copy from before your update, so what it cites is not what you published. Latency for a new page and for an updated page are different questions, and this study measures the first. The practical rule follows: never quote one crawl-to-citation figure for "AI search". Quote it per engine, with the date. Even inside Google the two surfaces disagree with each other, and the concordance study measures the spread.
Interpretation: how to read the number, and how not to
The intended use is expectation-setting. A site owner compares their own crawl-to-citation experience against a panel-derived benchmark. A latency far above the panel median is a concrete signal that something in the technical setup or authority profile is worth investigating — a better starting point than a vague feeling that AI citation "isn't happening yet". Four readings would be wrong.
"The median is N days, so my page should be cited by then." A median describes a population; half of it took longer. One page is a draw from a wide distribution, not a scheduled event.
"My page was crawled, so citation is coming." Crawling is necessary, not sufficient. Most crawled pages are never cited for any given query.
"Latency went down, so our optimisation worked." Or the engine changed, or the topics did. Without pages published over the same window as a comparison, those cannot be separated.
"Nothing in three weeks, so AI search doesn't work for us." Three weeks may be entirely normal. That sentence gets said so often precisely because no benchmark exists to check it against, and it is how good content programmes get cancelled early.
Null results we would publish
All three hypotheses can fail, and each failure publishes as prominently as a confirmation. Median latency exceeding fourteen days on every engine measured contradicts LT1. No relationship between domain authority and latency contradicts LT2. No meaningful difference between engines contradicts LT3. So does the outcome that would redirect the whole thread — a never-cited share so high that latency is not the interesting variable.
An infrastructure failure counts too: insufficient log quality across the panel to establish reliable start times. Saying "we could not measure this cleanly" is a real contribution in a field publishing numbers without saying how they got them. Each lands in the null results registry, under the standards this site publishes to.
Reproduction: measuring your own latency
A single-site version will not generalise, but it will tell you whether your own pipeline works at all, which is usually the real question. The citation playbook covers the page itself; the free tools here cover the mechanical parts.
- 1 Record exact publication timestamps Somewhere structured, not in your memory.
- 2 Set up durable, queryable log access If your host rotates logs weekly, ship them somewhere that keeps them longer than your study window.
- 3 Filter for retrieval bots and verify them Check status codes, and confirm the bot is what it claims where IP ranges are published. Record the first successful fetch per bot per URL.
- 4 Fix a small query set before you publish Three to five natural questions the page should answer. Written in advance, so you cannot fit them to the outcome.
- 5 Check on a sustainable schedule and log the negatives The absence of a citation on a given day is data, and dropping it is how the censoring problem starts.
- 6 Report both the distribution and the never-cited share Reporting only the first is the mistake this whole page warns against.
Who this applies to, and who it does not
A latency benchmark is most useful to publishers who ship new pages regularly and can instrument them. Frequency gives you a sample; instrumentation gives you a start time. Without both, the measurement is guesswork dressed as a metric.
It applies differently to a brand-new domain. With no crawl history, the publish-to-crawl stage may dominate everything else, so the crawl-to-citation figure is not the constraint that site is actually hitting — the discovery mechanisms above matter far more first. It applies weakly to sites whose value sits in a handful of stable pages: if you publish four pages a year, watching those four individually beats any distribution.
And it does not apply at all to content that will never be cited for reasons unrelated to timing. Gated material, thin pages and duplicated text have a citability problem that no amount of waiting solves, and reading a slow latency number as an engine problem in those cases points remediation at the wrong thing.
Raw data we will publish
Three files publish with the results. The per-page log extract carries verified first-fetch timestamps by bot. The daily citation-check log includes every negative check. The panel manifest describes each site's authority band, topic category and submission method, without naming sites that asked not to be named. Summary curves are never published without the rows behind them.
Limitations
- Sample: volunteers, not a random sample of the web. A recruited panel skews toward site owners already engaged enough with AI search to join a research panel, and toward those with the technical maturity to retain logs — both plausibly correlated with faster citation.
- Geography and language are uncontrolled. Panel sites will not be evenly distributed across regions or languages, and crawl schedules and index refreshes need not be uniform across them. Any result describes the panel's geography, not the web's.
- Engine coverage is limited to surfaces we can query daily. Engines without a checkable citation list, and newer agentic fetch modes, are outside the design entirely.
- Query selection bounds the outcome. "First citation" means first citation for a tracked query. A page citable for something outside the fixed set registers as never cited, which inflates the never-cited share by an unknown amount.
- Measurement resolution is one day. Daily checking cannot resolve sub-day latency, misses citations that appear and disappear between checks, and inherits any clock drift between log and monitor.
- Confounders survive the design. Page quality, topic competitiveness and publication timing are recorded but not randomised, so no causal claim about what shortens latency can come out of this study.
- Reproducibility depends on log quality we do not control. CDN edge caching can hide fetches from origin logs entirely, so some start times will be wrong in a direction we cannot detect.
- Generalisation is time-bounded. A latency figure describes the engines as they behaved during one collection window, and re-running is the only way to know whether it still holds.
The curves and the raw log files go out to the newsletter when the first cohort's window closes. If you publish regularly and keep durable server logs, the panel is still being recruited — there is no enrolment form, so the about page explains how to get in touch. The study sits inside the AI Citation Index programme.
Namdev, R. (2026). Crawl-to-citation latency (v1). Retrieved from https://ritiknamdev.com/blog/crawl-to-citation-latency-study Published under CC BY 4.0 — reuse freely with attribution.
Part of the AI Citation Index programme, which supplies the query set. Complements Zero to Cited, this site's own single-domain timeline, with a multi-site version; see also the citation half-life study, which asks how long a citation lasts once won.