None of these datasets exist yet. This page is the release plan for seven of them — what each will contain, the condition that has to be met before it ships, and the licence it ships under (CC BY 4.0, in every case). It is published in advance for one reason: a plan stated before the work is done is a plan someone can hold us to afterwards.
The release table: what ships, when, under what licence
| Dataset | Contents | Licence | Ships when | Status |
|---|---|---|---|---|
| Standard GEO Query Set measurement standard | One row per query: text, assigned category, intent classification, version it entered the set | CC BY 4.0 | Before any collection begins — so the sample can be challenged before results exist | Next up. Not live. |
| AI Bot User-Agent Registry the Registry | One row per bot: user-agent string, operator, category (training / retrieval / agentic), documented robots.txt policy, verification hostname pattern, date each field was last confirmed | CC BY 4.0 | Independent of collection — useful immediately, and improves with outside corrections | Planned. Not live. |
| robots.txt census raw data the Blocking Census | One row per domain × bot pair: domain, category, whether a robots.txt exists, the rule found for that bot, collection timestamp | CC BY 4.0 | Independent of collection | Planned. Not live. |
| AI Citation Index corpus the Citation Index | One row per citation: query text and category, engine, run number, timestamp, cited URL, cited domain, position in the response, context snippet | CC BY 4.0 | With the first full Index release — it cannot precede the collection it records | Planned. Not live. Largest file by a wide margin. |
| Query Fan-Out Corpus the Fan-Out study | One row per inferred sub-query: parent query, sub-query segment, classification against the patent's eight documented types, citations attributed to that segment | CC BY 4.0 | After the citation corpus has been published and checked — it is an inference layer on top of it | Planned. Not live. Carries the largest error bar. |
| Citation half-life tracking set the half-life study | One row per URL × engine × window: citation rate in that window, days since first observation | CC BY 4.0 | Rolling across a 12-month tracking period; interim windows published as they complete, labelled incomplete | Planned. Not live. |
| RCT panel results schema, freshness, author bio | One row per participating page: anonymised identifier, treatment or control assignment, before and after citation rate, collection windows, intervention applied | CC BY 4.0 | When each study concludes — dependent on volunteer recruitment reaching a workable size | Planned. Not live. Most exposed to slipping. |
Read the "ships when" column as a dependency graph rather than a calendar. No calendar dates are given here because none can be honestly promised; each release lands when the thing it rests on has been published and checked. Any slip gets posted on the relevant study page rather than quietly absorbed — a release plan that only ever reports on-time delivery is not a plan being tracked honestly.
- Seven datasets are planned. Zero are live. Treat every row above as a commitment, not an availability notice.
- All seven ship under CC BY 4.0 — commercial use permitted, attribution the only requirement, no share-alike clause.
- The query set ships before any results do. A query set published alongside favourable results invites the obvious suspicion that the queries were chosen to produce them.
- Corrections publish as new versions with a changelog entry. An already-published file is never silently edited, because a cited version has to stay stable.
- A plan is not a delivery. Judge this page against what actually ships, on the schedule each linked study states.
Why publish raw files instead of reports?
A chart or a percentage is a conclusion someone else drew from data you cannot see. A downloadable dataset is the thing itself: checkable, re-analysable, and reusable for a question the original report never asked. A summary is frozen at the moment its author wrote it; a dataset keeps answering questions that had not occurred to anyone when it was collected.
The gap being filled here is structural rather than a general preference for openness. Nearly every widely quoted AI-search number traces back to a visibility vendor. The best of that work is genuinely good: Ahrefs' AI SEO statistics roundup, Semrush's AI Overviews study, Profound's platform citation patterns and Zyppy's citation ranking factors are all carefully done. But for those companies the data is the product, so methods stay partially described and samples stay unpublished. Trace the most-quoted figures, as the provenance audit does, and the pattern is consistent: real sources, real numbers, no way for anyone outside to check them.
There is a live counter-example. When Originality.ai published its llms.txt tracking data, and when a 300,000-domain analysis found no clear citation effect, the method was concrete enough that does llms.txt actually work became answerable rather than contested. That is what a checkable dataset does to a debate. An independent publication has no product to protect: the credibility is the product, and credibility here means being checkable.
A chart is a conclusion someone else drew from data you can't see. A dataset is the thing itself — checkable, re-analyzable, reusable for a question the original report never asked.
Share on XWhat a row actually contains
Two fields in the citation corpus do most of the analytical work. Cited domain is what turns a pile of responses into an answer to the question of who actually gets cited. That is the question behind the most-cited domains analysis and narrower versions of it, like ChatGPT's dependency on Wikipedia. Those are claims that should be recomputable, not taken on trust. Run number matters more than it looks: because AI answers are non-deterministic, the same query across five runs produces the variance data that makes any single measurement interpretable at all.
In the bot registry, the load-bearing fields are category and last confirmed. Training crawlers, retrieval fetchers and agentic browsers behave nothing alike — the distinction the GPTBot versus OAI-SearchBot comparison turns on, and the one AI crawler statistics collapse whenever they report a single bot volume. A confirmation date is what separates a maintained registry from a stale listicle; compare against Cloudflare's 2025 crawler breakdown or Momentic's crawler reference and the value becomes obvious.
The path from a raw engine response to a file someone else can safely load is the same for every row in the table above, and it is the part most open-data promises skip over.
- 01 Collect Run the published query set against every engine, several runs per window, logging raw responses.
- 02 Structure Reduce responses to one row per citation, with the run number and timestamp preserved.
- 03 Document Write the data dictionary, the collection protocol, and the list of known quirks.
- 04 Privacy review Anonymise panel identity, drop raw logs, remove anything that could identify a person.
- 05 Version + publish Stamp a version, write the changelog entry, ship CSV and JSON under CC BY 4.0.
Why CC BY 4.0, and what versioning guarantees
The licence permits commercial use deliberately. A visibility vendor building a paid product on top of these files is an allowed and welcome outcome; a non-commercial clause would exclude exactly the organisations most able to extend the work. It requires attribution and nothing else, which is what keeps a citation chain traceable back to the original rows. It deliberately omits share-alike, so someone combining this with a proprietary dataset does not have to relicense their own work. And it is a licence people already recognise, which clears an organisation's legal review without a conversation.
The trade accepted is loss of control. Someone can take these files, build something better, and compete. If an openly published dataset produces better research than this site produces, the goal was met.
Every release carries an explicit version number. Corrections publish as a new version with a changelog entry stating what changed and why; the existing file is never silently edited, because a dataset that has already been cited needs a stable reference point.
What gets excluded, and what never does
Panel-dependent studies rest on site owners volunteering their pages, so their domains get anonymised — but the treatment assignment, the intervention and the outcome are all published. Raw server logs stay out because they contain IP addresses and request patterns that are not ours to publish regardless of consent; derived aggregates go out instead. Third-party page bodies stay out for copyright reasons and because full text adds nothing the URL and a snippet do not.
What never gets excluded is anything merely inconvenient. A finding that contradicts a claim made elsewhere on this site, or a data-quality problem discovered mid-collection, ships with the rest. Privacy is a genuine constraint; it is not a category that expands to cover embarrassment. Nulls are filed openly in the null results registry, and each dataset will appear against its study on the studies index, mirrored to a public repository once the collection tooling is released.
The bar a release has to clear
A file being downloadable is not the same as a dataset being usable. Publishing raw rows with no dictionary and no protocol technically satisfies "open data" while leaving it practically unusable, which is a common enough failure mode to be worth naming. A release missing any of the following is not ready, regardless of whether the underlying collection finished.
The worked example is the one that does the most work. A single documented analysis, with code, that reproduces a published figure from the raw rows makes a dataset usable faster than pages of prose documentation. It also changes what a published figure is: every number on the per-engine pages, from ChatGPT to Google AI Overviews, becomes a claim someone can falsify rather than a number they have to accept.
What publishing openly actually costs
Most arguments for open data skip the costs, which makes them less persuasive rather than more. Errors become permanent and public — a summary report with a mistake can be quietly corrected, a downloaded and cited dataset cannot, and versioning discipline is a mitigation rather than a fix. Preparation is real work: column definitions, formatting, quirk documentation and privacy review all take time that produces no new findings, and budgeting for it is the difference between a release plan and an intention.
Misuse is a certainty, not a risk. Someone will pull a row out of context or compute a ratio the sample cannot support. The accepted trade is that a verifiable number occasionally misused beats an unverifiable number universally trusted. It forecloses a business model — the obvious commercial path here is to build the measurement infrastructure and sell access, and publishing the raw output closes that door on purpose. And it invites scrutiny you cannot control, which is the entire point and still uncomfortable.
How to contribute
Two routes are open now: contribute queries to the Standard GEO Query Set if you have coverage gaps you want represented, or join the Citation Index panel if you would share anonymised server logs. Both are described on the Citation Index page.
A third costs less and helps more than it sounds — tell us where a planned schema in the table above is wrong. A field that turns out to be unusable is far cheaper to fix before collection than after. The people most likely to spot it are the ones who have already tried to answer a question like how a page actually earns a ChatGPT citation with the data currently available.
Nothing here is downloadable yet, so there is no file to fetch. Do the one thing that is useful today: check whether AI crawlers can reach your pages. Then subscribe to the newsletter — each release in the table above is announced there when it ships.
Limitations
- None of these datasets are live today. This page is the plan, not the download page.
- A plan is not a delivery. Everything here is a commitment made before the work is done. Judge it against what actually ships, on the schedule each linked study states.
- No calendar dates are given, because releases are gated on dependencies and on collection windows that have not run. A ship condition is stated instead; a date would be a guess.
- Panel-dependent datasets carry the most risk. The RCT results and log studies need volunteers. If recruitment falls short, those releases shrink or slip, and that constraint is outside our control.
- The fan-out corpus is inference, not observation. Sub-query boundaries are reconstructed from visible output because no public interface exposes the internal list. That dataset will carry a larger error bar than the others and should be read accordingly.
Namdev, R. (2026). The AI citation dataset strategy: what ships, when, and under what licence (v1). Retrieved from https://ritiknamdev.com/blog/ai-citation-dataset-strategy Published under CC BY 4.0 — reuse freely with attribution.
Every planned dataset ties to a study or reference asset already published — see the AI Citation Index for the flagship dataset this strategy centers on.