"Keep content fresh" appears on nearly every GEO tactic list, and no controlled study we could locate actually tests it against AI citation rate specifically. This page registers a randomized design: 300 stale pages, half genuinely updated, half left alone, citation rate measured before and after.
No data has been collected yet. Nothing on this page is a result. Participant sites are being enrolled. No page has been updated under the protocol.
- The tactic scoreboard grades content freshness Hypothesis — widely asserted, no controlled test located.
- 'Fresh' is ambiguous in practice: a changed timestamp and a substantive rewrite are different interventions.
- Whether updating a page changes its citation rate.
- Whether a timestamp change alone does anything, or whether substantive edits are required.
- How long any effect persists.
Why is the most universal GEO advice untested?
The research question: does a genuine content update causally raise a page's AI citation rate, holding topic and quality constant? Nobody knows, because the two obvious ways to find out both fail. Observational comparison confounds freshness with everything else that changes when a page is revised, and the word "fresh" is used so loosely that two people testing it would run two different experiments.
Content freshness is graded Hypothesis on the GEO tactic evidence scoreboard. It is widely recommended and mechanistically plausible: a more recently verified fact is a reasonable thing for a retrieval system to prefer, and the same reasoning appears in the original GEO paper. No controlled test isolates freshness from every other variable that changes when a page gets updated.
What does "fresh" actually mean here?
This is the ambiguity that has kept the question untested, so it is settled before anything else. Two different ideas hide behind the word. Recency is about when the page changed — a property of the file. Currency is about whether the content is still correct — a property of the world.
A page updated yesterday can be badly out of date; a page untouched for three years can be perfectly current if the subject has not moved. Advice that says "update your content" reaches for currency and measures recency, because recency is the part a machine can see. That gap is where date-stamp gaming comes from.
A genuine update for this study's purposes must meet at least one criterion, agreed with panel participants before enrollment. Adding at least one new, substantively informative section. Correcting a factual error or outdated figure. Or expanding existing coverage with new examples, data, or detail. Publishing genuinely new material is the version of this with the most independent support behind it. The argument that original research wins citations states it most clearly — and it is a different claim from the one under test here.
Simply changing a visible "last updated" date, reformatting without content change, or making only cosmetic edits does not qualify. A study testing only date-stamp changes would answer a narrower and less interesting question than the one this page asks.
Hypothesis One conceptual limitation of the outcome measure follows directly. If retrieval systems effectively register recency while readers care about currency, the two can diverge and a study measuring only citation rate would not notice. That is a limitation of what citation rate can tell us, not a flaw in the randomisation.
What circumstantial evidence exists?
Some circulating figures claim a high share of frequently-cited pages were recently updated — the freshness-and-AI-search coverage is a representative example. But these are observational, and they dead-end in the same way most numbers traced in where AI SEO statistics come from do. Recently-updated pages likely differ from stale ones in other ways: more actively maintained sites, more attention generally. A correlation between freshness and citation does not establish that updating a specific page causes its citation rate to rise.
Classic search has the same problem from the other end: Ahrefs' analysis of how old top-ranking pages are finds that high-ranking pages tend to be old, which is the opposite of what a naive freshness story predicts. Both patterns can be true at once, and neither is causal. The same structure appears in the brand-mentions correlation.
How is the study designed?
- 01 Recruit 300 pages Existing pages, stale 12+ months
- 02 Randomize Half get a genuine content update, half don't
- 03 Control the update type Substantive revision, not date-stamp only
- 04 Measure citation rate Before/after, treatment vs. control
- 05 Publish either result Pre-committed, positive or null
The registered hypothesis
H1: A genuine content update (new information, corrected facts, expanded coverage) to a stale page causally increases its AI citation rate within one collection cycle, relative to a matched control page left unchanged. It is graded Hypothesis until the data exists, on the same scale used in the visibility measurement standard and defined in the glossary.
'Keep content fresh' is on nearly every GEO tactic list. No controlled study we could find actually tests whether updating a page causes its AI citation rate to rise. We're registering the RCT that would.
Share on XWhich freshness signals actually exist?
Before asking whether freshness matters, it helps to list what could carry the signal. There are fewer candidates than the discussion usually implies.
| Signal | Where it lives | Why it may be weak |
|---|---|---|
| Visible date on the page | Page text | Trivially editable, so easy to discount |
| Structured date in markup | Page metadata | Frequently wrong or auto-generated |
| HTTP last-modified header | Server response | Often reflects deployment, not editing |
| Sitemap change date | Sitemap file | Commonly set by a template for every URL, which is why Google's own sitemap documentation treats it as a hint at best |
| Observed content change | The crawler's own comparison | Requires the crawler to have seen the page before — Google's crawler documentation and OpenAI's both describe separate fetchers with separate schedules |
| External signals | New links, new mentions | Not under the publisher's direct control |
Notice that the first four are all cheap to fake. A system that trusted them heavily would be easy to manipulate. That is a reason to suspect the last two matter more, though we have no evidence for the ordering.
Push-based notification is the one lever that shortens the first hop rather than the signal itself. IndexNow lets a publisher announce a change instead of waiting to be recrawled; Bing reports it speeding discovery, and the protocol itself is open and adopted by more than one engine. What none of that establishes is any effect on citation, which is a separate step further down the chain. Bing's own guidance on sitemaps in AI-powered search is explicit that discovery and selection are different problems — a distinction the Copilot guide goes into further.
What confound does the randomisation control?
This design randomly assigns which pages get updated, rather than comparing naturally-updated pages to naturally-stale ones. That breaks the link between "being the kind of site that updates content" and "getting updated content," isolating the update itself as the variable under test. It is the same design logic behind the schema markup RCT and the author E-E-A-T study, which share this panel.
Plausible mechanisms, in either direction
Suppose a positive effect exists. A plausible mechanism: a retrieval system weighting recency signals, whether through a last-modified date, a genuinely different crawl timestamp, or updated content simply displacing an outdated version in a training or retrieval corpus, would favor the freshly updated version. Which crawler sees the change first is not a detail. The fetch patterns described in the AI crawler statistics differ enough between bots that two engines can be looking at different versions of the same page on the same day.
Suppose no effect exists instead. A plausible explanation: current-generation retrieval leans more heavily on accumulated authority and existing citation patterns than on freshness signals specifically. Or any freshness signal genuinely present gets outweighed by other factors this design holds constant: topic, existing authority, backlink profile.
Can 300 pages detect the effect?
A randomised design is only useful if it can detect an effect of the size that would matter. That is a question about power, and it deserves stating before results exist.
Citation rate is a noisy outcome. The same page can be cited on one run and not the next — the run-to-run variance documented across repeated runs of the same query set is large enough to matter here. Noise of that kind eats power quickly.
Panel studies also lose participants. Sites redesign, get sold, or drop out. Attrition reduces the effective sample below whatever was recruited.
We will publish the power calculation and the minimum detectable effect alongside the results, rather than only afterwards. If the study turns out to be underpowered, that will be said in the same sentence as the finding.
An underpowered study that reports a null is not evidence of no effect. It is evidence that the study could not see one. Those get conflated constantly, and we would rather label it ourselves than have it labelled for us.
How could the measurement go wrong?
Measuring too soon. Propagation takes time and varies by engine — the crawl-to-citation latency study is the attempt to put a number on it. An early measurement biases toward the null.
Letting participants pick their own pages. People choose pages they already think will respond. Random assignment exists to prevent exactly that.
Changing more than the content. If an update also brings new internal links or a new template, the treatment is no longer just an update. A template change can alter rendering behaviour entirely, which is its own confound — the JavaScript rendering study covers why.
Ignoring the control group's drift. Control pages move too. Without them, ordinary variation looks like an effect.
Reporting only pages that were cited before. Pages that gain a first citation are part of the effect, and excluding them understates it.
Unblinded scoring. Whoever judges whether an update was genuine should not know the outcome. Otherwise the definition drifts toward whatever produced a result.
How far does an update have to travel?
Most arguments about freshness skip straight from "I changed the page" to "the answer should change." Between those two events sits a chain of steps, each with its own delay and its own failure mode. Naming them makes it obvious why measuring too early biases toward the null.
- 01 You publish the change The only step fully under your control
- 02 A crawler returns Recrawl interval varies by site, by engine, by page
- 03 The new version is stored Parsing, rendering and indexing, each with its own lag
- 04 Retrieval selects it Only for queries where the page is a candidate at all
- 05 An answer is regenerated And only then can any citation change be observed
Every hop is measurable by somebody, and almost none of it is measurable by you. Crawl frequency is the one place with public data: Cloudflare's breakdown of which bots crawl the web and its crawl-to-click analysis both show fetch volumes wildly out of proportion to referrals. Crawl-to-refer ratio work puts numbers on the same asymmetry. A high crawl rate means your change is seen quickly. It says nothing about whether it is used.
The indexing hop has its own well-documented lag. Onely's measurements of rendering delay describe a queue between fetch and usable content that a publisher never sees. And log-level accounts of newer fetchers, such as a thirty-day study of agentic crawler behaviour and analysis of how AI bots treat robots.txt, suggest that fetch behaviour is still changing month to month. Our own reading of that literature sits on the bot user-agent registry and the agentic browser page.
We do not know the length of any hop after the crawl for any engine. That ignorance is the reason the study waits a full collection cycle rather than measuring at a date chosen for convenience.
Running a smaller version yourself
You cannot randomise across three hundred sites. You can still do something much better than guessing.
1. Pick twenty comparable stale pages on your own site, using the fixed-query-set discipline described in the zero-to-cited log study. Similar topic, similar age, similar traffic.
2. Split them at random into ten and ten. Use a coin or a spreadsheet function. Do not choose by hand, because you will choose badly without meaning to.
3. Record a baseline. For a fixed set of questions, note whether each page is cited, over several runs, before you touch anything.
4. Update only the first ten. Make real improvements. Leave the other ten completely alone, however tempting it is.
5. Wait longer than feels reasonable. Six to eight weeks is a sensible minimum.
6. Re-measure both groups the same way. The control group is what tells you whether any movement is real.
7. Expect an inconclusive answer. With twenty pages you will probably not detect anything. That is worth learning directly, because it explains why the field has so little evidence.
When does this report?
Recruitment alongside the schema RCT panel, targeting the same Q1 2027 collection window as Index v1. The full sequence, and what is already fixed versus still open, is below.
- RegistrationNow
Design, hypothesis and analysis plan published before any data exists
This page is the registration document
- RecruitmentAlongside the schema RCT panel
Three hundred stale pages across participating sites
Revision criteria agreed with participants before enrollment
- BaselineBefore treatment
Repeated citation measurement on the fixed query set
Several runs, so that ordinary variance is visible
- Collection cycleQ1 2027
Treatment and control measured the same way, per engine
Same window as Index v1
- PublicationAfter analysis
Result, power calculation and raw data, whichever way it falls
A null is published with the same prominence as a positive
Null results we would publish
- No effect of a genuine update. H1 fails, and we publish it with the same prominence a positive result would get.
- A negative effect. If updated pages lose citations relative to controls, that is the most interesting outcome available and it would be reported without softening.
- An effect on one engine only. A split result is not a failure. It is a finding about architecture, and we would report it per engine.
- An underpowered study. If attrition leaves too few pages, we say the study could not answer the question rather than reporting a null as if it were evidence of absence.
Open questions
- Does the size of an update matter, or only that a change occurred?
- Is there a decay in any freshness benefit, and how quickly does it fade?
- Does updating a page change which passage gets quoted, even when the citation itself persists?
- Are time-sensitive topics affected differently from stable ones, as classic search suggests?
- Does removing outdated material help more than adding new material?
Limitations
- "Genuine update" still involves judgment about what counts as substantive. The exact revision criteria are published alongside panel recruitment to prevent post-hoc redefinition, but two editors could still classify the same edit differently.
- Update quality is not independently verified. Participants are asked for a good-faith improvement; a poorly executed revision could introduce errors or dilute focus, and the design cannot separate "was updated" from "was updated well."
- Sample and panel selection. Three hundred pages, drawn from sites willing to enrol in a randomised trial, are not a random sample of the web. Participating sites are likely better maintained than average, which is a plausible direction of bias.
- One collection cycle may be too short for an effect that takes longer to propagate through an engine's retrieval index — the propagation path above has at least five stages, each with its own lag. If the result is null this is flagged explicitly rather than read as an absence of effect.
- Engine coverage and generalisation. Results describe the engines measured in the collection window, in English, from one region. Recrawl behaviour differs sharply between engines, so a null on one engine is not a null on all of them, and an effect on one does not transfer.
- Confounds the randomisation does not remove. Randomisation balances page-level traits at assignment; it cannot control an engine changing its retrieval mid-cycle, a treated page attracting new external links because it was revised, or seasonality in the queries themselves.
- Reproducibility is partial. The protocol, criteria and analysis plan are public, but no third party can re-run this exact trial on these exact pages — replication means a new panel, and would be a genuinely different sample.
To find out whether your own updates are even being refetched: check which of your pages currently surface with the AI Overview Exposure Checker, then compare the dates against your server logs. To see how strongly freshness is evidenced against every competing tactic: read the grades on the GEO tactic evidence scoreboard.
No page has been updated under this protocol yet. The result, the power calculation and the raw data go out to the newsletter whichever way it falls. If you run a site with stale pages and would consider enrolling some of them, the about page explains how to get in touch — there is no enrolment form, the criteria get agreed in conversation before anything is randomised. The study runs inside the AI Citation Index programme, alongside the schema RCT and the citation half-life study, which asks the mirror-image question: how long a citation lasts once it exists.
Namdev, R. (2026). Does content freshness increase AI citations? (v1). Retrieved from https://ritiknamdev.com/blog/content-freshness-ai-citations-study Published under CC BY 4.0 — reuse freely with attribution.
Registered as one of the twelve untested hypotheses in the GEO tactic evidence scoreboard, and as Tier 3 research in the AI Citation Index.