Guide · Claude Code for SEO

Claude Code for SEO: the task list

Which SEO jobs an agent handles reliably, which it handles badly, and which should never be delegated — with the dry-run-first review discipline that makes the first category safe.

Ritik Namdev Ritik Namdev ·Published September 2026 ·12 min read ·Last verified September 2026

What an agent does reliably, and what it does not

The useful question is not "can an agent do SEO" but "which SEO jobs is the output checkable for". That is the boundary that holds. An agent is good at work you can verify against the codebase, and unreliable at judgements you cannot.

TaskReliabilityWhy
Parsing server logs for AI bot behaviourReliableLarge, repetitive, structured input — the task with the clearest advantage over a chat assistant
Finding missing or duplicate meta descriptionsReliableDeterministic check, output verifiable against the files
Mapping internal link structure, finding orphansReliableEnumeration, not judgement
Validating a sitemap against what is actually publishedReliableTwo lists compared; a mismatch is a fact
Drafting Article schema from values on the pageNeeds reviewEvery field must trace to real page content. A fabricated date passes a build cleanly
Proposing new internal linksNeeds reviewRoughly half of proposals are tenuous keyword matches rather than topical ones
Rewriting titles and descriptions in batchNeeds reviewFormulaic output is only visible when a batch is read side by side
Deciding which pages matter commerciallyUnreliableThe fact is not in the repository; asked anyway, it answers plausibly rather than declining
Judging whether a change is worth makingUnreliableIt can see a page lacks a description. It cannot know the page is being retired next month
Explaining why you are not citedUnreliableSource preference differs by engine and is not visible from your codebase at all
Measuring whether a change workedUnreliableNothing in this workflow observes outcomes. That needs instrumentation and time
What this guide establishes
  • The reliable tasks share one property: the output can be checked against the codebase. That is the whole selection rule.
  • The dry-run-first, human-approves-second workflow is what makes the middle category usable. Every serious failure mode passes a build cleanly.
  • Nothing here is measured against outcomes. This guide is about doing the work faster, not evidence that the work produces citations.
  • The binding constraint on a team is review quality, not model capability.
  • Cost figures in circulation, including the $5-15 per session estimate below, are borrowed from existing coverage. We have not published a first-party benchmark.

The safe/unsafe delegation table

Autonomy is a spectrum, not a switch. Here is where to set it, and what each row is protecting against. Concrete examples of what happens when these boundaries slip are catalogued in Claude Code SEO failure modes, which doubles as a review checklist for exactly this table.

CapabilitySensible defaultWhy
Reading files, running buildsAllow freelyNothing leaves the machine. The cost of asking every time exceeds the risk.
Editing local files on a branchAllow within a named scopeReviewable as a diff, revertible in one command.
Generating structured dataPropose only; review every fieldA fabricated date, author or statistic passes a build cleanly. This is the single failure this site treats as unacceptable.
Editing canonicals, robots rules, redirect maps, sitemapsExplicit approval, every timeThese fail invisibly on the rendered page and are expensive to unwind once crawled.
Pushing, deploying, publishingNever automaticThe only step outside the revert-and-retry model.
Anything touching production dataNever automaticSame reason, with worse consequences.

Note what the middle rows have in common. It is not that indexing directives are hard to write — they are trivial. It is that a wrong one produces a perfectly healthy-looking page while quietly removing it from an index, and nobody notices until traffic moves weeks later. Verification for those changes has to be external: fetch the live file, check the header, watch the logs. A passing build proves nothing about them, and the GPTBot versus OAI-SearchBot split is the clearest case of a one-line change with a large invisible consequence.

What Claude Code actually is, for SEO purposes

A CLI-based agentic coding tool — documented by Anthropic, with a separate web-search tool for fetching pages — that reads a codebase, proposes changes, runs builds, and with permission applies changes directly. For SEO work that means it operates on a site's actual source files: templates, meta tags, schema, sitemaps. Not advice a human then implements by hand.

That is a difference in kind from a chat assistant, not just convenience, for any task spanning dozens or hundreds of pages. It also sits on the publisher side of the wider shift toward an agent-readable web: the agent here reads your source because your site declares nothing an agent could call. It is also different from a traditional SEO crawler, which enumerates problems deterministically and then stops. The strongest setup uses both: the crawler as the source of truth about what is broken, the agent to propose and apply the specific fix. It is the method behind the zero-to-cited log study and the llms.txt test on this site.

The dry-run-first workflow

Dry-run-first audit pipeline
  1. 01 Crawl the site Read pages, sitemap, robots.txt
  2. 02 Identify issues Missing schema, broken links, thin meta
  3. 03 Propose changes Dry-run diff, nothing applied yet
  4. 04 Human review Approve, reject, or edit each change
  5. 05 Apply + verify Build passes, spot-check in browser

The critical discipline is step 3: propose changes as a reviewable diff before anything touches the live site. Alongside it, five practices do most of the safety work.

  • Run the build after every change, not just at the end of a session — an early error is cheaper than debugging a large batch.
  • Spot-check visually in a real browser for anything touching layout. A passing build is a syntax result, not a visual one.
  • Commit incrementally so any single bad change is easy to isolate and revert.
  • Never let it invent a value. A fabricated date, author or statistic passes a build cleanly. The provenance audit exists because so much published SEO material fails this test.
  • Codify the rules in a file. Anything that must hold every time — design tokens, tone, what must never be invented — belongs in an instruction file the tool reads on start. A rule stated once in conversation does not survive the session.

Claude Code for SEO isn't about generating marketing copy with a prompt. It's an execution agent operating directly on a site's codebase — audits, schema, structural fixes — with a dry-run-first, human-approves-second discipline.

Share on X

Prompt patterns that work and fail

The difference between a useful session and a mess is usually the instruction, not the model.

Weak instructionStronger versionWhy it matters
"Improve the SEO on this site""List pages under this directory with a missing meta description"Names the scope and the check. Produces a reviewable list, not edits.
"Add schema to the blog""For this one post, draft Article schema using only values present on the page"Constrains the source of every field. Makes fabrication visible.
"Fix the internal linking""Propose up to three contextual links per page, and show the sentence each would sit in"Caps the change and forces the reasoning into the output.
"Make the titles better""Rewrite these ten titles. Show old and new side by side. Change nothing yet"A batch shown together is the only way to catch formulaic output.
"Clean up the redirects""Read the current redirect map and report any chains or loops"Diagnosis before treatment on a file where mistakes are expensive.

Three habits produce most of the improvement: name the scope, ask for output before edits, cap the size of the change. A fourth is less obvious and just as useful — ask the agent what it is unsure about before it proceeds. An admission of uncertainty is far cheaper to handle than a confident wrong value.

The first hour, step by step

The first hour, in the order that keeps mistakes cheap
  1. 01 Work on a branch Everything else assumes a bad change is one command from disappearing.
  2. 02 Ask for a description "Describe how meta descriptions are generated here" — tests understanding before any edit.
  3. 03 Check it against what you know Subtly wrong here means subtly wrong in a diff later. The cheapest reliability test there is.
  4. 04 One change, one file Read the diff line by line. Apply. Build. Open the page.
  5. 05 Write down the project rules Tokens, tone, what must never be invented. An instruction file is how rules survive the session.
  6. 06 Only now scale up Directory, then page type. Increase scope after each successful review pass, not before.

Step 3 is the one people skip and the one that pays. If the agent's description of your codebase is subtly wrong before it has edited anything, every diff it later proposes inherits that error. Checking a description costs a minute and tells you more about reliability than any benchmark.

A worked example: reading server logs

This is the task with the clearest advantage over a chat assistant, because a log file is large, repetitive and structured — exactly what an agent handles well.

The instruction shape. Point at the log file. Ask for a table of user-agent by request count by status code, the top requested paths per bot, and anything returning an error. Match agent strings against the published references — OpenAI's, Perplexity's, Anthropic's and Google's, or Momentic's consolidated list, cross-checked against our own bot registry.

What to check first in the output. Error responses to crawlers — a problem you can fix today, with no theory required. Then whether a bot you believe is blocked is in fact still requesting, and whether training crawlers and retrieval bots are behaving differently.

The trap. User-agent strings are self-reported and can be spoofed. A line claiming to be a major crawler may not be one; verification requires a reverse lookup on the requesting address, and an agent will not do that unless you ask.

The limit. Logs tell you what was fetched, not what was used. A crawl is not a citation, and the gap between them is the whole open question in this field — taken up in the crawl-to-citation latency study. For the shape of output worth aiming at, Paul Calvano's robots.txt and AI bots analysis is the reference, and Google's crawl-budget documentation is how to interpret request volume.

Context management on a large site

The limiting factor on real sites is not capability but how much the agent can hold in view at once.

Work in readable unitsOne template, one directory, one page type. An instruction spanning a large site gets executed on the part it happened to read.
Fresh session per unitA long session accumulates context, and early instructions compete with everything since.
Rules in a file, not a messageA rule stated once in conversation does not survive the session.
Log what each session changedCommit messages work. Six sessions later, this is how you answer "when did that change, and why".

One more constraint worth naming: do not let it optimise for a metric it cannot see. Performance is the clearest case — Core Web Vitals are field measurements with documented thresholds and a population baseline, and an agent editing code observes none of them.

What it costs, and what drives the bill

Open question

Existing coverage estimates roughly $5-15 in API credits for a session auditing hundreds of pages and proposing metadata and internal-linking fixes. That is a rough, source-dependent figure, not a controlled benchmark. We have not published a first-party cost benchmark, so treat any circulating figure as an anecdote from one site with one codebase.

What the bill responds to is more useful than a number. How much gets read, not how much gets written — an audit that reads five hundred pages to change ten is dominated by the reading, so narrowing scope is the largest single lever. Session length, because a long session re-processes accumulated context repeatedly; several short scoped sessions generally cost less than one long one. Rework, because a vague instruction that produces an unusable first pass costs twice. And file structure: large monolithic templates cost more to work with than small focused ones, for the same reason they are harder for people.

Who this workflow suits, and who it does not

Honestly, a narrower group than most coverage suggests.

  • Fits: people with file-level access and version control — static sites, templated CMS setups with source access, headless stacks. Without both, the review-and-revert discipline this depends on does not exist.
  • Fits: repetitive, structural work across many similar pages, where the leverage is real and review cost stays manageable.
  • Fits: people who can read a diff. Every safety practice here assumes a competent reviewer.
  • Fits badly: a handful of pages. Setup and review overhead can exceed the work itself.
  • Does not fit: closed no-code platforms without file access, which would need a different integration entirely.
  • Does not help: a site whose content is only assembled by client-side JavaScript. No amount of agent-authored markup changes what a retrieval bot can read.

Four misreadings lead somewhere bad, and they are worth naming. "So I can automate my SEO" — you can automate execution; diagnosis, prioritisation and judgement remain yours. "The dry run is optional once you trust it" — trust is not the variable, since the failures that matter are the ones nobody notices. "A passing build means the change is good" — a build is a syntax check, and every serious failure mode in the companion page passes one cleanly. "An agent-written page will get cited" — nothing here supports that, and none of the citation behaviour measured on ChatGPT or Perplexity turns on how a page was authored.

Working this way in a team

A team changes the risk profile in both directions. The good part is that review already exists: a team with mandatory code review has the single most important safeguard built into its process, and most of the discipline argued for above is simply their existing workflow.

The bad part is that reviews get lazier as volume rises. An agent produces far more change than a human author would, and a reviewer facing a forty-file diff behaves differently from one facing a four-file diff. Throughput can quietly outrun attention, and nothing in the tooling will tell you when it does.

  • Label the origin of agent-assisted commits. Later investigation then narrows fast.
  • Agree the autonomy boundary as a team, and write it where the tool can read it. Otherwise the risk level is set by the most permissive person.
  • Keep one person accountable per session.
Hypothesis

Our expectation is that review quality, not model capability, is the binding constraint on how safely a team can use this workflow. That is reasoning from the failure modes we have catalogued, not a measured finding.

Limitations

  • Nothing here is measured against outcomes. Everything is about doing the work faster. Nothing shows the work produces citations, and a study connecting the two would change what this guide should recommend.
  • No first-party cost benchmark exists. The estimate above is borrowed.
  • Vertical differences are out of scope — local, commerce and YMYL sites each carry constraints this general workflow does not address.
  • This is a workflow description, not a guaranteed cost or outcome for any specific site.
Where to go next

Take the one task off the agent that it fails silently: generate your structured data in the schema generator, where every field is one you typed, then have the agent place it. A fabricated date or author inside JSON-LD passes a build cleanly, which is why the delegation table above marks that row "propose only".

Then read the review checklist that makes the rest of the table enforceable: the failure modes catalogue lists what has actually gone wrong. The reusable-component version of this workflow is covered in open-source SEO agent tooling.

How to cite this
Namdev, R. (2026). Claude Code for SEO: the task list (v1). Retrieved from https://ritiknamdev.com/blog/claude-code-for-seo-guide

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

See Claude Code SEO failure modes for what goes wrong, and open-source SEO agent tooling for the reusable component approach.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

Anthropic — Claude documentationdocs.anthropic.com Anthropic — Web search tool documentationplatform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool Anthropic Support — Does Anthropic crawl the web, and how do I block it?support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler OpenAI — Bots and crawlers documentationplatform.openai.com/docs/bots OpenAI developer docs — bots referencedevelopers.openai.com/api/docs/bots Perplexity — crawler referencedocs.perplexity.ai/docs/resources/perplexity-crawlers Google Search Central — Build and submit a sitemapdevelopers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap Google Search Central — Overview of Google crawlers and fetchersdevelopers.google.com/search/docs/crawling-indexing/overview-google-crawlers Google Search Central — Managing crawl budget for large sitesdevelopers.google.com/search/docs/crawling-indexing/large-site-managing-crawl-budget Google Search Central — Core Web Vitalsdevelopers.google.com/search/docs/appearance/core-web-vitals Google Search Central — AI optimization guidedevelopers.google.com/search/docs/fundamentals/ai-optimization-guide web.dev — Defining the Core Web Vitals thresholdsweb.dev/articles/defining-core-web-vitals-thresholds HTTP Archive Web Almanac 2025 — Performance chapteralmanac.httparchive.org/en/2025/performance llmstxt.org — the llms.txt proposalllmstxt.org Ahrefs — What is llms.txt?ahrefs.com/blog/what-is-llms-txt Ahrefs — llms.txt study across 300k domainsahrefs.com/blog/llmstxt-study Search Engine Journal — Google says llms.txt is purely speculative for nowwww.searchenginejournal.com/google-says-llms-txt-is-purely-speculative-for-now/577576 Ahrefs — Schema markup and AI citationsahrefs.com/blog/schema-ai-citations Paul Calvano — AI bots and robots.txtpaulcalvano.com/2025-08-21-ai-bots-and-robots-txt Technology Checker — robots.txt AI crawler blocking reporttechnologychecker.io/blog/robots-txt-ai-crawlers-blocking-report Momentic — AI search crawlers and bots referencemomenticmarketing.com/blog/ai-search-crawlers-bots Bing — IndexNow getting startedwww.bing.com/indexnow/getstarted Wikipedia — IndexNowen.wikipedia.org/wiki/IndexNow
FAQ

Frequently asked questions

Is this the same as generic AI content writing?
No. This is about using Claude Code as an execution agent for technical SEO tasks — audits, schema generation, structural fixes, log analysis — that operate on a codebase. It is not about generating marketing copy.
How much does a typical session cost?
Figures circulating for a session auditing hundreds of pages sit in the low tens of dollars of API credits, but we could not trace any of them to a disclosed method or a named source, and this site has not measured its own. Costs scale with site size and task complexity, so treat this as a rough starting point, not a fixed benchmark. We have not published a first-party benchmark.
What is the biggest risk of using an agent this way?
Applying changes without review. The dry-run-first, human-approves-second workflow exists specifically to prevent an agent pushing an unreviewed change to a live site. See the failure-modes companion page for concrete examples of what goes wrong when that discipline is skipped.
Does this require a codebase — can it help a site built on a no-code platform?
The workflow assumes direct file access: a static site, a templated CMS with accessible source, or a headless setup. A fully no-code, closed platform without file-level access would need a different integration, likely through that platform's own API rather than direct codebase edits.
How large a site can this realistically handle in one session?
Smaller than most people expect, and the limit is context rather than capability. The practical pattern is to work one directory, template or page type at a time, with a fresh session for each. A large site is many small sessions, not one big one.
Do I need to be a developer to use this?
You need enough familiarity to read a diff and know when it looks wrong. The dangerous position is being able to run the tool but not evaluate the output, because every failure mode depends on a reviewer catching something. If you cannot review the change, you are not supervising the agent, you are trusting it.
Can it replace an SEO crawler tool?
Not sensibly. A crawler enumerates issues deterministically across a whole site, cheaply and repeatably. An agent is better at proposing and applying a specific fix. The strongest setup runs the crawler to find the problems, then hands a scoped list to the agent to fix, keeping the deterministic tool as the source of truth about what is broken.
What should I never let it do automatically?
Anything that reaches production without a human decision, and anything that changes indexing directives. Deploys, pushes, canonical tags, robots rules, redirect maps and sitemap regeneration all belong behind an explicit approval — not because they are hard, but because their failures are invisible on the rendered page and expensive to unwind once crawled.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.