Nothing has shipped yet
There is no repository, no download, no release date and no waitlist. This page describes software that does not exist publicly. It is published in advance so the design and the commitment are on the record before the code, rather than after — the same ordering this site applies to study predictions. If you came here looking for a tool to install, stop reading and go to the tools that actually exist instead.
- A stated intention: when each major study publishes, the matching collection code gets published alongside it in a public repository.
- A design sketch of what that code contains, so someone can build their own version now rather than waiting.
- An argument that the record format matters more than the code — a shared schema would do more for this field than any single release.
- An honest cost list, because the reason so few independent datasets exist is that people underestimate the maintenance, not the build.
- No repository is live. No date has been set. Nothing on this page should be read as an announcement.
Why open-source the tooling at all
Nearly every visibility vendor's methodology is proprietary, because their business model depends on it being so. An independent research publication has the opposite incentive: reproducibility is the credibility, not a cost against it.
Every dataset published here — the Citation Index, the crawler statistics, the provenance audit — is only as trustworthy as its collection method. The strongest proof a method is sound is someone else running it and getting a comparable result, and that requires the code, not a description of it.
It costs little competitively. Anyone can run the same scripts; nobody can replicate a year of longitudinal data, a recruited panel of participating sites, or a track record of non-retracted findings. Meaningful divergence from a reproduction would itself be useful: either it exposes an unstated assumption in the original method, or an execution difference on the reproducer's side. Both beat an unreproducible number nobody can check — the situation most vendor statistics in this field are stuck in.
Every visibility vendor's methodology is proprietary because their business model depends on it. An independent research publication has the opposite incentive — reproducibility is the credibility, not a cost against it.
Share on XWhat is planned for release
Each of these automates work an agent can already get wrong in a specific way; the failure catalogue records which.
| Tool | What it would do | Status |
|---|---|---|
| Citation collection harness | Runs the fixed query set against tracked engines, extracts and structures citation records | Not released |
| Server-log analysis pipeline | Parses logs for AI bot classification, crawl-to-referral ratios, crawl-to-citation latency | Not released |
| Technical audit agents | Reusable Claude Code skills for schema checks, meta-description audits, internal-link mapping | Not released |
| robots.txt census crawler | The collection code behind the robots.txt AI-blocking census | Not released |
Bare scripts would not be enough. A release that is reproducible in practice rather than in theory needs three more things. Documented environment requirements, with exact versions. A stated collection protocol covering query timing, session state and retry logic. And example output showing the expected shape of results. Code without those three is technically open and practically unreproducible.
Two limits are worth naming now. The audit agents can check facts deterministically — is a crawler blocked, does the page render server-side, is the structured data valid, are canonicals coherent. They cannot tell you whether you will be cited, because nobody has established the causal factors, and any tool outputting a confident visibility score is inventing a quantity.
The genuinely hard parts
If someone set out to build this today, almost none of the time would go where a plan predicts.
- Defining a citation. One engine shows a numbered list, another inline markers, another names a publisher in prose with no link. Deciding what counts and applying it consistently is the central methodological problem, and code cannot solve it.
- Handling variance. The same question asked twice returns different sources, so every collection has to repeat queries — multiplying cost and complicating every downstream count.
- Keeping conditions constant. Region, session state, device and time all plausibly affect results. Holding them fixed across quarters is discipline, not engineering, and easy to break by accident.
- Surviving interface changes. Adapters break. The question is whether they break loudly, with a failed run, or quietly, with subtly wrong parsing that corrupts a quarter of data first.
- Verification in the log pipeline. A user-agent string is self-reported and easy to forge; confirming a request came from who it claims requires checking published address ranges. Cloudflare has documented crawlers that did not identify honestly. Any crawler-traffic figure that skipped this step counted impostors too.
- Storage discipline. Keeping raw responses is boring, and it is what lets you answer next year's question. Almost everyone learns this after discarding them.
Agent frameworks changed one half of this and not the other. Reading a messy answer page and extracting structured records without a hand-written parser got dramatically easier, and so did auditing thousands of pages for whether they answer early or attribute their statistics. Deciding what counts as a citation, holding conditions constant, and knowing whether a number is true are not parsing problems, so better parsing does not touch them. One thing got harder: a model-based extractor may behave slightly differently between runs or versions, which has to be pinned and recorded, or the tooling introduces exactly the variance it exists to measure.
The record format matters more than the code
The most reusable output of this project may not be software at all. If several independent groups collect citation data in the same shape, results can be pooled and compared. If everyone invents their own shape, every dataset is an island — which is roughly where the field is now.
A shared record would carry, at minimum, the query and panel version, the engine and surface, the timestamp, the collection conditions, the full source list, whether each source was named in the answer text, and a reference to the stored raw response. Most of that is context rather than result. That ratio is the point: a citation record without its conditions cannot be compared with anything. The dataset strategy sets out what each release would contain.
Our expectation is that a shared schema would do more for this field than any single study. We cannot demonstrate that, and it is labelled as an opinion rather than a finding.
What running this would cost you
Open tooling is free. Running it is not, and pretending otherwise sets people up to abandon a programme halfway. Collection time: a panel of a thousand queries across several engines, repeated for variance, is real hours or real API spend. Storage: raw responses only accumulate, and need a home that outlives whoever set it up. Maintenance: adapters break on someone else's schedule, usually just before a quarterly run. Judgement: the largest cost and the least visible, because someone has to decide edge cases consistently and write down what they decided, or the series drifts.
Releasing code adds obligations a research page does not have — issues arrive, breakage becomes public, forks diverge and report different numbers under a similar name, and documentation rots. The mitigation for the forking problem is precisely published conditions, not restrictive licensing. Collection also touches other people's systems, so the defaults are where good behaviour lives. Rate limits on by default rather than as an option someone has to find. Honest identification. Respect for stated preferences. And query panels containing questions rather than people.
Build your own version now, without waiting
You do not need this release. A small version is genuinely achievable, and the small version is where most of the practical value sits.
- Start with a spreadsheet, not a codebase. Twenty questions, three engines, columns for linked and named, collected by hand. You will learn more about the measurement problems in one afternoon than in a week of building.
- Write your citation definition down first, before collecting anything. You will be tempted to adjust it later to make the numbers nicer; the written version is what stops you.
- Keep the raw answers. Paste them somewhere. They cost nothing to store and are impossible to recover.
- Automate only what hurts. Log parsing is worth scripting early because it is genuinely repetitive. Answer collection is worth keeping manual longer than you would expect, because that is where you notice what a script would silently discard.
- Version your query panel. Note every addition and removal. A panel that changes silently makes every previous quarter incomparable.
Who this is for, and who it is not
Research infrastructure has a narrower audience than a tool. It is for other researchers wanting to check a published figure or run the same method elsewhere. For in-house technical teams that want their own numbers rather than a vendor average — the group most likely to get value quickly. And for agencies with a developer, for whom running a fixed panel per client is a genuinely differentiated service. It is not, initially, for non-technical site owners. Making it usable without a terminal is a worthwhile goal and a separate project, and pretending otherwise would waste their time.
Limitations
- No repository is live — this page describes an intention and a design, not a shipped product, and no release date has been set.
- Whether open tooling shifts a field's norms is an open question. Open question It has worked elsewhere. It may not work here. Publishing anyway is the only way to find out.
- Nothing here is a measurement. No claim on this page rests on data.
Run the first column of your own panel today, with no code: the AI Overview Exposure Checker tells you which of your queries produce a generated answer at all, which is the population any citation panel has to be drawn from. Record it, date it, and you have step one of the spreadsheet method above.
Then automate the one audit check that is genuinely deterministic: the schema generator produces the valid JSON-LD the audit agents described here would only verify. For the collection method these scripts would implement, written out in full, the llms.txt log study is the closest published example.
Namdev, R. (2026). Open-source SEO agent tooling: what to build (v1). Retrieved from https://ritiknamdev.com/blog/open-source-seo-agent-tooling Published under CC BY 4.0 — reuse freely with attribution.
Complements Claude Code for SEO and supports the AI Citation Index's reproducibility commitment.