Commitment · Nothing released yet

Open-source SEO agent tooling: what to build

No code has been released. This page states, before the fact, what the citation-collection harness and log-analysis pipeline behind this site's research would contain — and why the record format matters more than the code.

Ritik Namdev Ritik Namdev ·Published September 2026 ·8 min read ·Last verified September 2026

Nothing has shipped yet

There is no repository, no download, no release date and no waitlist. This page describes software that does not exist publicly. It is published in advance so the design and the commitment are on the record before the code, rather than after — the same ordering this site applies to study predictions. If you came here looking for a tool to install, stop reading and go to the tools that actually exist instead.

What this page is
  • A stated intention: when each major study publishes, the matching collection code gets published alongside it in a public repository.
  • A design sketch of what that code contains, so someone can build their own version now rather than waiting.
  • An argument that the record format matters more than the code — a shared schema would do more for this field than any single release.
  • An honest cost list, because the reason so few independent datasets exist is that people underestimate the maintenance, not the build.
  • No repository is live. No date has been set. Nothing on this page should be read as an announcement.

Why open-source the tooling at all

Fact

Nearly every visibility vendor's methodology is proprietary, because their business model depends on it being so. An independent research publication has the opposite incentive: reproducibility is the credibility, not a cost against it.

Every dataset published here — the Citation Index, the crawler statistics, the provenance audit — is only as trustworthy as its collection method. The strongest proof a method is sound is someone else running it and getting a comparable result, and that requires the code, not a description of it.

It costs little competitively. Anyone can run the same scripts; nobody can replicate a year of longitudinal data, a recruited panel of participating sites, or a track record of non-retracted findings. Meaningful divergence from a reproduction would itself be useful: either it exposes an unstated assumption in the original method, or an execution difference on the reproducer's side. Both beat an unreproducible number nobody can check — the situation most vendor statistics in this field are stuck in.

Every visibility vendor's methodology is proprietary because their business model depends on it. An independent research publication has the opposite incentive — reproducibility is the credibility, not a cost against it.

Share on X

What is planned for release

Each of these automates work an agent can already get wrong in a specific way; the failure catalogue records which.

ToolWhat it would doStatus
Citation collection harnessRuns the fixed query set against tracked engines, extracts and structures citation recordsNot released
Server-log analysis pipelineParses logs for AI bot classification, crawl-to-referral ratios, crawl-to-citation latencyNot released
Technical audit agentsReusable Claude Code skills for schema checks, meta-description audits, internal-link mappingNot released
robots.txt census crawlerThe collection code behind the robots.txt AI-blocking censusNot released

Bare scripts would not be enough. A release that is reproducible in practice rather than in theory needs three more things. Documented environment requirements, with exact versions. A stated collection protocol covering query timing, session state and retry logic. And example output showing the expected shape of results. Code without those three is technically open and practically unreproducible.

Two limits are worth naming now. The audit agents can check facts deterministically — is a crawler blocked, does the page render server-side, is the structured data valid, are canonicals coherent. They cannot tell you whether you will be cited, because nobody has established the causal factors, and any tool outputting a confident visibility score is inventing a quantity.

The genuinely hard parts

If someone set out to build this today, almost none of the time would go where a plan predicts.

  • Defining a citation. One engine shows a numbered list, another inline markers, another names a publisher in prose with no link. Deciding what counts and applying it consistently is the central methodological problem, and code cannot solve it.
  • Handling variance. The same question asked twice returns different sources, so every collection has to repeat queries — multiplying cost and complicating every downstream count.
  • Keeping conditions constant. Region, session state, device and time all plausibly affect results. Holding them fixed across quarters is discipline, not engineering, and easy to break by accident.
  • Surviving interface changes. Adapters break. The question is whether they break loudly, with a failed run, or quietly, with subtly wrong parsing that corrupts a quarter of data first.
  • Verification in the log pipeline. A user-agent string is self-reported and easy to forge; confirming a request came from who it claims requires checking published address ranges. Cloudflare has documented crawlers that did not identify honestly. Any crawler-traffic figure that skipped this step counted impostors too.
  • Storage discipline. Keeping raw responses is boring, and it is what lets you answer next year's question. Almost everyone learns this after discarding them.

Agent frameworks changed one half of this and not the other. Reading a messy answer page and extracting structured records without a hand-written parser got dramatically easier, and so did auditing thousands of pages for whether they answer early or attribute their statistics. Deciding what counts as a citation, holding conditions constant, and knowing whether a number is true are not parsing problems, so better parsing does not touch them. One thing got harder: a model-based extractor may behave slightly differently between runs or versions, which has to be pinned and recorded, or the tooling introduces exactly the variance it exists to measure.

The record format matters more than the code

The most reusable output of this project may not be software at all. If several independent groups collect citation data in the same shape, results can be pooled and compared. If everyone invents their own shape, every dataset is an island — which is roughly where the field is now.

A shared record would carry, at minimum, the query and panel version, the engine and surface, the timestamp, the collection conditions, the full source list, whether each source was named in the answer text, and a reference to the stored raw response. Most of that is context rather than result. That ratio is the point: a citation record without its conditions cannot be compared with anything. The dataset strategy sets out what each release would contain.

Hypothesis

Our expectation is that a shared schema would do more for this field than any single study. We cannot demonstrate that, and it is labelled as an opinion rather than a finding.

What running this would cost you

Open tooling is free. Running it is not, and pretending otherwise sets people up to abandon a programme halfway. Collection time: a panel of a thousand queries across several engines, repeated for variance, is real hours or real API spend. Storage: raw responses only accumulate, and need a home that outlives whoever set it up. Maintenance: adapters break on someone else's schedule, usually just before a quarterly run. Judgement: the largest cost and the least visible, because someone has to decide edge cases consistently and write down what they decided, or the series drifts.

Releasing code adds obligations a research page does not have — issues arrive, breakage becomes public, forks diverge and report different numbers under a similar name, and documentation rots. The mitigation for the forking problem is precisely published conditions, not restrictive licensing. Collection also touches other people's systems, so the defaults are where good behaviour lives. Rate limits on by default rather than as an option someone has to find. Honest identification. Respect for stated preferences. And query panels containing questions rather than people.

Build your own version now, without waiting

You do not need this release. A small version is genuinely achievable, and the small version is where most of the practical value sits.

  1. Start with a spreadsheet, not a codebase. Twenty questions, three engines, columns for linked and named, collected by hand. You will learn more about the measurement problems in one afternoon than in a week of building.
  2. Write your citation definition down first, before collecting anything. You will be tempted to adjust it later to make the numbers nicer; the written version is what stops you.
  3. Keep the raw answers. Paste them somewhere. They cost nothing to store and are impossible to recover.
  4. Automate only what hurts. Log parsing is worth scripting early because it is genuinely repetitive. Answer collection is worth keeping manual longer than you would expect, because that is where you notice what a script would silently discard.
  5. Version your query panel. Note every addition and removal. A panel that changes silently makes every previous quarter incomparable.

Who this is for, and who it is not

Research infrastructure has a narrower audience than a tool. It is for other researchers wanting to check a published figure or run the same method elsewhere. For in-house technical teams that want their own numbers rather than a vendor average — the group most likely to get value quickly. And for agencies with a developer, for whom running a fixed panel per client is a genuinely differentiated service. It is not, initially, for non-technical site owners. Making it usable without a terminal is a worthwhile goal and a separate project, and pretending otherwise would waste their time.

Limitations

  • No repository is live — this page describes an intention and a design, not a shipped product, and no release date has been set.
  • Whether open tooling shifts a field's norms is an open question. Open question It has worked elsewhere. It may not work here. Publishing anyway is the only way to find out.
  • Nothing here is a measurement. No claim on this page rests on data.
Where to go next

Run the first column of your own panel today, with no code: the AI Overview Exposure Checker tells you which of your queries produce a generated answer at all, which is the population any citation panel has to be drawn from. Record it, date it, and you have step one of the spreadsheet method above.

Then automate the one audit check that is genuinely deterministic: the schema generator produces the valid JSON-LD the audit agents described here would only verify. For the collection method these scripts would implement, written out in full, the llms.txt log study is the closest published example.

How to cite this
Namdev, R. (2026). Open-source SEO agent tooling: what to build (v1). Retrieved from https://ritiknamdev.com/blog/open-source-seo-agent-tooling

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

Complements Claude Code for SEO and supports the AI Citation Index's reproducibility commitment.

FAQ

Frequently asked questions

Where can I download it?
You cannot. No repository is live and no release date has been set. This page describes an intention and a design, published so the commitment is on the record before the code exists rather than after.
Is there a waitlist?
No. There is no signup form for this specifically. Releases would be announced alongside the study they support.
Doesn't open-sourcing the tooling give away the whole advantage?
No. The value in the research programs on this site is the longitudinal data collected over time, not the collection scripts. Releasing the code costs little competitively and buys reproducibility, which is the scarce resource in this field.
Is this the same as the AI Bot Registry or the Citation Index datasets?
Related, but distinct. The Registry and Index are data and reference assets. This is about the reusable code that would produce and maintain them.
Can I trust results produced by someone else running this code?
Only as far as their disclosed collection conditions go. Identical code with a different region, date, account state or query panel produces genuinely different numbers. Code makes a claim checkable; it does not make it correct.
What is the single hardest part to get right?
Normalising what counts as a citation across engines that present sources completely differently. Everything else is engineering. That one is a judgement call that has to be written down and applied consistently, or nothing compares to anything.
Does open tooling risk flooding engines with automated queries?
It is a fair concern, which is why rate limits, sane defaults and honest identification belong in the code rather than in a README nobody reads. Research access should behave like a considerate guest, and the defaults are where that gets enforced.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.