Citation Half-Life is the number of days until a cited URL's citation rate falls to half its peak value. It requires tracking the same set of URLs across repeated collection windows over a full year — patience, not cleverness. Nothing like it exists in public data today.
No data has been collected yet. Nothing on this page is a result. Collection runs across four quarterly checkpoints beginning Q1 2027. No citation has been tracked yet.
- Citation sets change composition over roughly a year — Ahrefs measured its AI Overview top-10 overlap figure twice, about a year apart, and it fell substantially.
- Every published AI-citation study located for this page reports a snapshot of a population, not a cohort followed forward.
- The methodological precedent exists: fix a subject, observe it repeatedly over a defined window, publish the window.
- Whether an individual cited URL keeps its citation, and for how long.
- Whether decay follows a smooth curve, a sharp cliff, or some other shape entirely.
- Whether half-life differs by engine, or by whether a page is actively updated.
Why no one knows how long a citation lasts
Every published AI-citation study reports a snapshot: what was cited, on a given date, for a given query set. None of them, as far as we could verify, re-check the same set of citations weeks or months later to see whether they're still there. The graded inventory on the statistics hub makes the pattern visible in one place, and the reason it recurs is traced in where AI SEO statistics actually come from.
That gap has teeth. A citation that disappears after two weeks and one that persists for a year are scored identically by every existing measurement, yet they are entirely different outcomes for whoever earned them. Almost every piece of practical advice in the field — collected and graded on the tactic evidence scoreboard — quietly assumes a citation earned is a citation held. Nobody has checked.
The financial consequence is the reason this is worth a year of patience. If citations decay quickly, content spend is an operating cost that recurs forever and "getting cited" is a weaker goal than "staying cited". If they are durable, content spend is capital expenditure that depreciates slowly and initial placement dominates. Those are different businesses, and nobody currently knows which one they are in.
What "citation half-life" means here
Citation Half-Life is the number of days until a URL's citation rate across the tracked query set falls to 50% of its peak observed rate. A URL cited in 80% of relevant runs at peak and dropping to 40% has reached its half-life at that point, regardless of whether it disappears entirely afterward.
The physics term is borrowed deliberately. Radioactive decay, drug metabolism and post-engagement decay share a shape: a quantity starting at a peak and falling away, well described by how long it takes to halve. Using the word commits to a testable claim. Not the unfalsifiable "citations fade eventually", but "citations fade in a comparable, quantifiable pattern across URLs and engines" — which this study can confirm or refute. The surrounding vocabulary is defined in the glossary, and the framing itself belongs to GEO before it belongs to statistics.
What would a decay curve look like?
If real data eventually resembled this illustrative shape, the half-life point would fall between months 5 and 6, where the curve crosses 50 on the vertical axis. Whether real decay is this smooth, faster, slower, or a sharp cliff rather than a curve is what the study is designed to establish.
What snapshot studies can and cannot tell us
A few published measurements come close enough to durability to be worth reading carefully, and seeing exactly where each stops is the clearest way to understand what a longitudinal design adds.
The nearest thing to a trend claim is Ahrefs' measurement of how many AI Overview citations come from the top ten results, run twice roughly a year apart and showing a substantial fall, as Search Engine Journal covered at the time. That is two snapshots of a population, not one cohort followed forward. It tells you the composition of the citation set changed; it cannot tell you whether any individual URL held its citation, because the URLs were never tracked as individuals. The full context sits on the AI Overview statistics page.
The important distinction is between a population and a cohort. A population whose composition is stable can be made entirely of URLs that churn constantly; a population whose composition shifts can be made of URLs that mostly stayed put. Only cohort tracking separates those two worlds. Every published design we reviewed — Semrush's AI Overviews study, Ahrefs' overlap work, Profound's platform citation patterns — is a population design.
The precedent worth borrowing outright comes from a different question entirely. A thirty-day site-log study of agentic crawler behaviour demonstrates the shape: fix a subject, observe it repeatedly over a defined window, publish the window. Nothing about that is difficult. It is slower than a snapshot, which is why almost nobody does it.
Older fields did the same work under different names — academic link rot, digital preservation, and the practitioner notion of content decay, whose closest published measurement is Ahrefs' work on how long ranking takes and how old top-ranking pages are. All of them found decay to be normal, measurable, and mostly explained by maintenance and competition. Whether AI citation behaves the same way — or whether engine changes dominate in a way those fields never faced — is the open question.
Evidence Citation sets change composition over periods of about a year. That is supported. Open question Whether individual citations persist, and for how long, is not measured anywhere we could find.
Study design and schedule
- Month 1–3Q1 2027
Baseline citation snapshot from Index v1.
Establishes which URLs are cited at time zero.
- Month 4–6Q2 2027
First re-check of the same URL set.
Early decay pattern becomes visible.
- Month 7–9Q3 2027
Second re-check.
First half-life estimates become statistically meaningful.
- Month 10–12Q4 2027
Full-year survival curve.
Published alongside the State of AI Search annual report.
Tracking begins with the same URL set identified in Index v1's first collection, re-checked at months 4, 7, and 10. First meaningful half-life estimates publish in Q3 2027; the full 12-month survival curve lands in the State of AI Search annual report.
Pre-registered hypotheses
| # | Hypothesis | Predicted outcome |
|---|---|---|
| HL1 | A majority of cited URLs remain cited 90 days later | Expect not to hold |
| HL2 | Half-life varies significantly by engine — some engines refresh their citation set faster than others | Expect to hold |
| HL3 | Pages that are actively updated (per their own metadata) show longer half-life than static pages | Expect to hold |
HL1's predicted direction, against persistence, is deliberately the more provocative bet. Registering it up front means nobody can accuse us of shaping the result after seeing it — and if most citations turn out to be durable for 90+ days, that is a genuinely reassuring finding.
HL3 is the hypothesis most likely to be misread. It predicts that pages reporting active updates hold citations longer. It does not predict that updating a page causes it to hold a citation longer; that is a causal claim requiring a controlled design, registered separately as the content freshness study. The recency preference this trades on is widely asserted and almost never tested — the same criticism that applies to Zyppy's correlational citation ranking factors and to the effect sizes in the peer-reviewed GEO benchmark study, which is the strongest evidence the field has and still measures a snapshot.
Every AI citation study reports a snapshot. None track the same citations forward. We're registering a prediction that most citations do NOT survive 90 days — precisely so a surprising result can't be dismissed as us shaping the finding.
Share on XFour routes by which a citation is lost
"The citation went away" is one observation with at least four causes underneath it. Naming them separately is what turns a number into an explanation.
Open question Which route dominates is unknown. It is entirely possible that route three accounts for most observed loss, in which case almost all publisher-side advice about maintenance is beside the point.
Route one has an upstream dependency worth naming: a change to your page only becomes a change to the engine's view of your page once a crawler returns. Recrawl intervals are not uniform, as Google's own crawl-budget documentation makes clear, and rendering adds further lag on top of fetching, per Onely's rendering-delay measurements. Whether AI retrieval bots render at all is an open question with its own study, and the end-to-end lag is registered as the crawl-to-citation latency study.
Route three has a harder dependency, because access to it belongs to the operators. What little is documented publicly sits in vendor documentation: Anthropic's web search tool, Perplexity's developer docs and Google's documentation of AI features in Search. None of them states how often source selection is recomputed, which is exactly the number that would let route three be separated from route two without guessing.
Why two identical half-lives can mean opposite things
The following figures are invented to illustrate why one number is not enough. No measurement is being reported here.
Site A starts with a citation on a reference page. The rate slides gradually, month by month, reaching half after roughly six months. Nothing dramatic happens on any single date. Site B holds a steady rate for five months, then falls off almost entirely within a fortnight and stays low. Its half-life also lands around six months.
Both would report the same headline figure, and the situations are opposite. Site A looks like ordinary competitive erosion, answered by steady maintenance. Site B looks like an event — a model update, a policy change, one strong competitor arriving — answered by finding out what happened on that date rather than by rewriting the page.
Hypothesis We expect both patterns to appear in real data, which is why the study reports curves rather than a single average. If only one pattern occurs, that is itself a finding worth stating. Reading those curves takes a few habits.
Edge cases the definition has to survive
| Edge case | Ruling | Why it matters |
|---|---|---|
| Never cited | Excluded | No peak means no half-life. Counting it as instant decay would fabricate an observation. |
| Cited once, never again | Depends on fixed run count | A single observation may be a citation or noise; only a fixed number of runs makes the two distinguishable. |
| URL redirects | Continuity, and recorded | A followed redirect to substantially the same content is the same citation, but the fact of the move is logged. |
| Domain cited via a different page | Recorded on both axes | The site kept the citation; the URL did not. Reporting only one of those misleads. |
| Citation returns | Non-monotonic, tracked as such | A metric assuming one-way decline will misdescribe every URL that comes back. |
Every one of these decisions has to be fixed before data collection, not after. Choosing them afterwards is how an honest study drifts into a flattering one. The general form of that commitment, applied across this site's research programme, is set out in the measurement standard.
What will not transfer between engines
| Dimension | Why it breaks transfer |
|---|---|
| Refresh cadence | Live retrieval at query time can change an answer daily; a periodically rebuilt index cannot. |
| Source-set size | An engine naming three sources per answer churns more visibly than one naming ten. |
| Recency preference | Where new material is explicitly favoured, decay is partly designed behaviour rather than competitive outcome. |
That these differences exist is not speculation. Comparative work on how each assistant chooses sources and Ziptie's account of ChatGPT's source selection both describe engines behaving differently at the selection stage. How far any two of them agree is the subject of the concordance study.
Source composition compounds it. An engine leaning on encyclopedic reference material has a structurally different decay profile from one leaning on community discussion, because those corpora change at different rates — compare the Wikipedia dependency in ChatGPT against the Reddit dependency in Perplexity. Surface matters too: AI Mode and AI Overviews draw on different sources, so a half-life attributed to "Google" without naming the surface is already ambiguous.
Because of this, the study reports per engine. A blended half-life across engines would be a number with no referent in the world.
Confounds and measurement pitfalls
A study that runs for a year accumulates confounds simply by existing. Model updates mean any result describes a moving system, not a fixed one. Query drift makes a fixed query set less representative over time. Our own influence — publishing about a query set can in principle change what is written about those topics; the effect is probably tiny, not zero. Sampling attrition: sites disappear, get redesigned, or block crawlers mid-study. Seasonality: annual cycles look like decay if the window straddles a peak.
The pitfalls on the analyst's side are separate and mostly self-inflicted. Confusing absence with loss: an engine answering a different kind of question has not dropped your citation. Changing the query wording, which makes a new experiment rather than a re-check. Single runs, where between-run variation can exceed the effect you are looking for. Editing while measuring, which resets the baseline. Tracking only winners, which says nothing about pages that were never cited.
How to track citation persistence yourself
You do not need a finished half-life number to hedge sensibly: treat every citation you earn as something to maintain rather than a one-time achievement. You can also run a small version of this study today. It is tedious rather than difficult.
- 1 Pick twenty questions Where your pages would be a reasonable answer. Write them down exactly.
- 2 Run each three times Per engine, in a clean session. Record which sources were named.
- 3 Note the date Plus anything unusual: a model release, a site change, a news event.
- 4 Do nothing for a month Without a stable period you cannot separate decay from your own edits.
- 5 Recheck at 1, 3 and 6 months Uneven spacing reveals shape better than evenly spaced checks.
- 6 Guess the route on each loss Your guess may be wrong. It forces the right question.
- 7 Report a range Twenty questions and three runs is an impression, not a measurement.
The prior question, for most sites, is presence rather than persistence — there is nothing to maintain if nothing is cited yet. The getting-cited guide covers that ground. The cheapest hedge once you are cited is to keep a small set of pages untouched as a control while you refresh the rest. Without one, a citation that returns after a refresh gets credited to the refresh whether or not it earned it.
Common misreadings of a half-life number
| The misreading | What the number actually supports |
|---|---|
| Content is worthless after the half-life | It is a median. A large share of citations last considerably longer. |
| Short half-life means refresh everything | Only if route one dominates. Engine-driven loss is unaffected by refreshing. |
| Long half-life means publish and forget | A durable citation on a page that is now wrong is a liability, not an asset. |
| It measures content quality | It measures selection behaviour over time, shaped by much more than the page. |
The same caution applies when reporting decay upward. A half-life is not a planning constant. It is a median with a wide spread, measured on a system that changes, so name the date on any figure you quote. The useful framing for a stakeholder is whether the company is buying a durable asset or renting attention — the answer decides whether content spend is capital or operating expense. It matters most to sites whose value rests on a few pages being cited repeatedly, and least to news and time-bound content where decay was always expected.
Null results we would publish
- No decay at all. If citation rates are flat across a year, the metric is unnecessary and we will retire it publicly.
- No engine difference. HL2 predicts variation between engines. If they behave alike, the prediction fails and we report it as a failure.
- No update effect. If actively updated pages decay at the same rate as static ones, HL3 fails and a great deal of common advice loses its main support.
- No usable curve. If the data is too noisy to fit any shape, we will publish the noise rather than smoothing it into a story.
Three things would change this page outright:
- Another group publishing a longitudinal citation dataset — which we would prefer to being first.
- An operator documenting how often its source selection is recomputed.
- Evidence that citation loss is mostly non-monotonic, with sources cycling in and out. That would break the decay framing entirely.
Where this sits in the research programme
This study is one of several registered against the studies index, and deliberately the slowest. Everything else can produce a result inside a collection window; this one cannot produce anything meaningful before a year has passed, which is precisely why nobody has run it. The baseline URL set comes from the Index's first collection, whose sampling is described in the dataset strategy, and the query set is shaped by the query fan-out corpus study, since a question that decomposes into many sub-queries offers more routes to citation and more routes to lose it.
Four questions the design will not settle on its own:
- Whether a citation lost to one engine tends to be lost to the others at the same time.
- Whether a survivor effect exists — a citation that lasts six months being much more likely to last a year.
- Whether domains have a half-life distinct from their individual URLs.
- Whether decay runs faster for questions with many acceptable answers than for questions with one correct answer.
Whatever this study finds, including nothing, goes to the null results registry and into the annual state of AI search summary. The methodology commitments behind that promise are in the method notes.
Limitations
- A URL's disappearance from citation could reflect many causes — content decay, competitive displacement, or model updates — which this study is not initially designed to distinguish between.
- Twelve months is a long baseline for a field this fast-moving; the underlying engines may change meaningfully mid-study, which will be logged as a confound rather than hidden.
- The baseline query set excludes several categories that have their own reasons to behave differently — local search, commerce and YMYL topics — so no finding here extends to them.
To establish your own time-zero baseline before decay can be measured at all: check which of your URLs are currently surfaced with the AI Overview Exposure Checker, and record the date. To follow the collection this study depends on: the schedule and sampling frame live in the AI Citation Index.
Namdev, R. (2026). AI Citation Half-Life (v1). Retrieved from https://ritiknamdev.com/blog/ai-citation-half-life-study Published under CC BY 4.0 — reuse freely with attribution.
Registered as one of the four novel metrics defined for the AI Citation Index.