Dataset strategy · Roadmap

The AI citation dataset strategy: what ships, when, and under what licence

Seven planned open datasets, in one table — contents, licence, ship condition and status. Nothing here is live yet; this is the release plan, published in advance so it can be checked against what actually ships.

Ritik Namdev Ritik Namdev ·Published September 2026 ·Planned, not yet live ·10 min read ·Last verified September 2026
The short version

None of these datasets exist yet. This page is the release plan for seven of them — what each will contain, the condition that has to be met before it ships, and the licence it ships under (CC BY 4.0, in every case). It is published in advance for one reason: a plan stated before the work is done is a plan someone can hold us to afterwards.

The release table: what ships, when, under what licence

DatasetContentsLicenceShips whenStatus
Standard GEO Query Set
measurement standard
One row per query: text, assigned category, intent classification, version it entered the setCC BY 4.0Before any collection begins — so the sample can be challenged before results existNext up. Not live.
AI Bot User-Agent Registry
the Registry
One row per bot: user-agent string, operator, category (training / retrieval / agentic), documented robots.txt policy, verification hostname pattern, date each field was last confirmedCC BY 4.0Independent of collection — useful immediately, and improves with outside correctionsPlanned. Not live.
robots.txt census raw data
the Blocking Census
One row per domain × bot pair: domain, category, whether a robots.txt exists, the rule found for that bot, collection timestampCC BY 4.0Independent of collectionPlanned. Not live.
AI Citation Index corpus
the Citation Index
One row per citation: query text and category, engine, run number, timestamp, cited URL, cited domain, position in the response, context snippetCC BY 4.0With the first full Index release — it cannot precede the collection it recordsPlanned. Not live. Largest file by a wide margin.
Query Fan-Out Corpus
the Fan-Out study
One row per inferred sub-query: parent query, sub-query segment, classification against the patent's eight documented types, citations attributed to that segmentCC BY 4.0After the citation corpus has been published and checked — it is an inference layer on top of itPlanned. Not live. Carries the largest error bar.
Citation half-life tracking set
the half-life study
One row per URL × engine × window: citation rate in that window, days since first observationCC BY 4.0Rolling across a 12-month tracking period; interim windows published as they complete, labelled incompletePlanned. Not live.
RCT panel results
schema, freshness, author bio
One row per participating page: anonymised identifier, treatment or control assignment, before and after citation rate, collection windows, intervention appliedCC BY 4.0When each study concludes — dependent on volunteer recruitment reaching a workable sizePlanned. Not live. Most exposed to slipping.

Read the "ships when" column as a dependency graph rather than a calendar. No calendar dates are given here because none can be honestly promised; each release lands when the thing it rests on has been published and checked. Any slip gets posted on the relevant study page rather than quietly absorbed — a release plan that only ever reports on-time delivery is not a plan being tracked honestly.

What this page establishes
  • Seven datasets are planned. Zero are live. Treat every row above as a commitment, not an availability notice.
  • All seven ship under CC BY 4.0 — commercial use permitted, attribution the only requirement, no share-alike clause.
  • The query set ships before any results do. A query set published alongside favourable results invites the obvious suspicion that the queries were chosen to produce them.
  • Corrections publish as new versions with a changelog entry. An already-published file is never silently edited, because a cited version has to stay stable.
  • A plan is not a delivery. Judge this page against what actually ships, on the schedule each linked study states.

Why publish raw files instead of reports?

A chart or a percentage is a conclusion someone else drew from data you cannot see. A downloadable dataset is the thing itself: checkable, re-analysable, and reusable for a question the original report never asked. A summary is frozen at the moment its author wrote it; a dataset keeps answering questions that had not occurred to anyone when it was collected.

The gap being filled here is structural rather than a general preference for openness. Nearly every widely quoted AI-search number traces back to a visibility vendor. The best of that work is genuinely good: Ahrefs' AI SEO statistics roundup, Semrush's AI Overviews study, Profound's platform citation patterns and Zyppy's citation ranking factors are all carefully done. But for those companies the data is the product, so methods stay partially described and samples stay unpublished. Trace the most-quoted figures, as the provenance audit does, and the pattern is consistent: real sources, real numbers, no way for anyone outside to check them.

There is a live counter-example. When Originality.ai published its llms.txt tracking data, and when a 300,000-domain analysis found no clear citation effect, the method was concrete enough that does llms.txt actually work became answerable rather than contested. That is what a checkable dataset does to a debate. An independent publication has no product to protect: the credibility is the product, and credibility here means being checkable.

A chart is a conclusion someone else drew from data you can't see. A dataset is the thing itself — checkable, re-analyzable, reusable for a question the original report never asked.

Share on X

What a row actually contains

Two fields in the citation corpus do most of the analytical work. Cited domain is what turns a pile of responses into an answer to the question of who actually gets cited. That is the question behind the most-cited domains analysis and narrower versions of it, like ChatGPT's dependency on Wikipedia. Those are claims that should be recomputable, not taken on trust. Run number matters more than it looks: because AI answers are non-deterministic, the same query across five runs produces the variance data that makes any single measurement interpretable at all.

In the bot registry, the load-bearing fields are category and last confirmed. Training crawlers, retrieval fetchers and agentic browsers behave nothing alike — the distinction the GPTBot versus OAI-SearchBot comparison turns on, and the one AI crawler statistics collapse whenever they report a single bot volume. A confirmation date is what separates a maintained registry from a stale listicle; compare against Cloudflare's 2025 crawler breakdown or Momentic's crawler reference and the value becomes obvious.

The path from a raw engine response to a file someone else can safely load is the same for every row in the table above, and it is the part most open-data promises skip over.

From raw response to published dataset
  1. 01 Collect Run the published query set against every engine, several runs per window, logging raw responses.
  2. 02 Structure Reduce responses to one row per citation, with the run number and timestamp preserved.
  3. 03 Document Write the data dictionary, the collection protocol, and the list of known quirks.
  4. 04 Privacy review Anonymise panel identity, drop raw logs, remove anything that could identify a person.
  5. 05 Version + publish Stamp a version, write the changelog entry, ship CSV and JSON under CC BY 4.0.

Why CC BY 4.0, and what versioning guarantees

The licence permits commercial use deliberately. A visibility vendor building a paid product on top of these files is an allowed and welcome outcome; a non-commercial clause would exclude exactly the organisations most able to extend the work. It requires attribution and nothing else, which is what keeps a citation chain traceable back to the original rows. It deliberately omits share-alike, so someone combining this with a proprietary dataset does not have to relicense their own work. And it is a licence people already recognise, which clears an organisation's legal review without a conversation.

The trade accepted is loss of control. Someone can take these files, build something better, and compete. If an openly published dataset produces better research than this site produces, the goal was met.

Every release carries an explicit version number. Corrections publish as a new version with a changelog entry stating what changed and why; the existing file is never silently edited, because a dataset that has already been cited needs a stable reference point.

What gets excluded, and what never does

Panel participant identityDomains anonymised. Assignment, intervention, and outcome all published.
Raw server logsDerived aggregates only — raw logs carry IPs that are not ours to publish.
Anything identifying a personExcluded rather than published and hoped over.
Third-party page bodiesURLs, domains, and short context snippets — never full text.
Not excluded: anything inconvenientContradictory findings and data quality problems ship with the rest.

Panel-dependent studies rest on site owners volunteering their pages, so their domains get anonymised — but the treatment assignment, the intervention and the outcome are all published. Raw server logs stay out because they contain IP addresses and request patterns that are not ours to publish regardless of consent; derived aggregates go out instead. Third-party page bodies stay out for copyright reasons and because full text adds nothing the URL and a snippet do not.

What never gets excluded is anything merely inconvenient. A finding that contradicts a claim made elsewhere on this site, or a data-quality problem discovered mid-collection, ships with the rest. Privacy is a genuine constraint; it is not a category that expands to cover embarrassment. Nulls are filed openly in the null results registry, and each dataset will appear against its study on the studies index, mirrored to a public repository once the collection tooling is released.

The bar a release has to clear

A file being downloadable is not the same as a dataset being usable. Publishing raw rows with no dictionary and no protocol technically satisfies "open data" while leaving it practically unusable, which is a common enough failure mode to be worth naming. A release missing any of the following is not ready, regardless of whether the underlying collection finished.

Data dictionaryEvery column defined: type, permitted values, and what a null means.
Collection protocolQuery timing, session state, geography, and retry logic stated in full.
Known quirks disclosedMid-run engine changes, empty responses, rows affected by a bug found later.
Row count and file sizeSo you can confirm a complete download and pick the right tool.
A worked exampleOne documented analysis, with code, reproducing a published figure from raw rows.

The worked example is the one that does the most work. A single documented analysis, with code, that reproduces a published figure from the raw rows makes a dataset usable faster than pages of prose documentation. It also changes what a published figure is: every number on the per-engine pages, from ChatGPT to Google AI Overviews, becomes a claim someone can falsify rather than a number they have to accept.

What publishing openly actually costs

Most arguments for open data skip the costs, which makes them less persuasive rather than more. Errors become permanent and public — a summary report with a mistake can be quietly corrected, a downloaded and cited dataset cannot, and versioning discipline is a mitigation rather than a fix. Preparation is real work: column definitions, formatting, quirk documentation and privacy review all take time that produces no new findings, and budgeting for it is the difference between a release plan and an intention.

Misuse is a certainty, not a risk. Someone will pull a row out of context or compute a ratio the sample cannot support. The accepted trade is that a verifiable number occasionally misused beats an unverifiable number universally trusted. It forecloses a business model — the obvious commercial path here is to build the measurement infrastructure and sell access, and publishing the raw output closes that door on purpose. And it invites scrutiny you cannot control, which is the entire point and still uncomfortable.

How to contribute

Two routes are open now: contribute queries to the Standard GEO Query Set if you have coverage gaps you want represented, or join the Citation Index panel if you would share anonymised server logs. Both are described on the Citation Index page.

A third costs less and helps more than it sounds — tell us where a planned schema in the table above is wrong. A field that turns out to be unusable is far cheaper to fix before collection than after. The people most likely to spot it are the ones who have already tried to answer a question like how a page actually earns a ChatGPT citation with the data currently available.

Next step

Nothing here is downloadable yet, so there is no file to fetch. Do the one thing that is useful today: check whether AI crawlers can reach your pages. Then subscribe to the newsletter — each release in the table above is announced there when it ships.

Limitations

  • None of these datasets are live today. This page is the plan, not the download page.
  • A plan is not a delivery. Everything here is a commitment made before the work is done. Judge it against what actually ships, on the schedule each linked study states.
  • No calendar dates are given, because releases are gated on dependencies and on collection windows that have not run. A ship condition is stated instead; a date would be a guess.
  • Panel-dependent datasets carry the most risk. The RCT results and log studies need volunteers. If recruitment falls short, those releases shrink or slip, and that constraint is outside our control.
  • The fan-out corpus is inference, not observation. Sub-query boundaries are reconstructed from visible output because no public interface exposes the internal list. That dataset will carry a larger error bar than the others and should be read accordingly.
How to cite this
Namdev, R. (2026). The AI citation dataset strategy: what ships, when, and under what licence (v1). Retrieved from https://ritiknamdev.com/blog/ai-citation-dataset-strategy

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

Every planned dataset ties to a study or reference asset already published — see the AI Citation Index for the flagship dataset this strategy centers on.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

arXiv - GEO: Generative Engine Optimization (the paper that named the field)arxiv.org/abs/2311.09735 Ahrefs - AI SEO statistics roundupahrefs.com/blog/ai-seo-statistics Semrush - AI Overviews studywww.semrush.com/blog/semrush-ai-overviews-study Profound - AI platform citation patternswww.tryprofound.com/blog/ai-platform-citation-patterns Zyppy - AI citation ranking factorssignal.zyppy.com/p/ai-citation-ranking-factors Ziptie - How original research wins AI citationsziptie.dev/blog/how-original-research-wins-ai-citations Google Patents - US11663201B2, query fan-outpatents.google.com/patent/US11663201B2 Search Engine Journal - Query fan-out technique in AI Modewww.searchenginejournal.com/query-fan-out-technique-in-ai-mode-new-details-from-google/552532 Cloudflare - From Googlebot to GPTBot: who is crawling your site in 2025blog.cloudflare.com/from-googlebot-to-gptbot-whos-crawling-your-site-in-2025 Momentic - AI search crawlers and bots referencemomenticmarketing.com/blog/ai-search-crawlers-bots Paul Calvano - AI bots and robots.txtpaulcalvano.com/2025-08-21-ai-bots-and-robots-txt Technology Checker - robots.txt AI crawler blocking reporttechnologychecker.io/blog/robots-txt-ai-crawlers-blocking-report BuzzStream - Study of publishers blocking AI crawlerswww.buzzstream.com/blog/publishers-block-ai-study Ahrefs - Overlap between AI search enginesahrefs.com/blog/ai-search-overlap Seer Interactive - 87% of SearchGPT citations match Bing top resultswww.seerinteractive.com/insights/87-percent-of-searchgpt-citations-match-bings-top-results Originality.ai - llms.txt tracking studyoriginality.ai/blog/llms-txt-tracking-study Search Engine Journal - llms.txt shows no clear effect across 300k domainswww.searchenginejournal.com/llms-txt-shows-no-clear-effect-on-ai-citations-based-on-300k-domains/561542 Discovered Labs - How ChatGPT, Claude and Perplexity choose sourcesdiscoveredlabs.com/blog/ai-citation-patterns-how-chatgpt-claude-and-perplexity-choose-sources Wikipedia - Generative engine optimizationen.wikipedia.org/wiki/Generative_engine_optimization Google Search Central - AI features and your websitedevelopers.google.com/search/docs/fundamentals/ai-optimization-guide Ahrefs - Schema markup and AI citationsahrefs.com/blog/schema-ai-citations
FAQ

Frequently asked questions

Are any of these datasets available today?
No. This page describes the release plan, tied to each study's own schedule. Check the linked study pages for specific timelines. None of these datasets are live yet.
Why release raw data instead of just the summary report?
Because you can't independently re-check a summary report. This whole site's reproducibility argument depends on someone else being able to load the real rows and verify a finding themselves.
Will there be a cost to access any of this?
No — the intent throughout is CC BY 4.0, free to reuse with attribution, consistent with every dataset commitment made elsewhere on this site.
What happens if an error is discovered in a dataset after release?
A correction gets published as a new version. The changelog entry states exactly what changed and why. We never silently edit the existing file. That would break anyone relying on a specific cited version.
Could a business use these datasets commercially, including a paid product built on top of them?
Yes. CC BY 4.0 explicitly permits commercial use. Attribution is the only requirement. Building a tool or service on top of an openly licensed dataset is exactly what this license is meant to enable.
How big will these files be?
The citation corpus is the largest by a wide margin: roughly 1,000 queries times seven engines times five runs per collection window produces tens of thousands of rows per release. The registry and query set are small enough to read in a text editor. Sizes get stated on each release.
Why publish a release plan before anything has shipped?
For the same reason the studies are pre-registered. A plan published in advance is checkable later. If a dataset slips or gets quietly dropped, that is visible against this page. A plan published only after delivery cannot be held to anything.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.