What an agent does reliably, and what it does not
The useful question is not "can an agent do SEO" but "which SEO jobs is the output checkable for". That is the boundary that holds. An agent is good at work you can verify against the codebase, and unreliable at judgements you cannot.
| Task | Reliability | Why |
|---|---|---|
| Parsing server logs for AI bot behaviour | Reliable | Large, repetitive, structured input — the task with the clearest advantage over a chat assistant |
| Finding missing or duplicate meta descriptions | Reliable | Deterministic check, output verifiable against the files |
| Mapping internal link structure, finding orphans | Reliable | Enumeration, not judgement |
| Validating a sitemap against what is actually published | Reliable | Two lists compared; a mismatch is a fact |
Drafting Article schema from values on the page | Needs review | Every field must trace to real page content. A fabricated date passes a build cleanly |
| Proposing new internal links | Needs review | Roughly half of proposals are tenuous keyword matches rather than topical ones |
| Rewriting titles and descriptions in batch | Needs review | Formulaic output is only visible when a batch is read side by side |
| Deciding which pages matter commercially | Unreliable | The fact is not in the repository; asked anyway, it answers plausibly rather than declining |
| Judging whether a change is worth making | Unreliable | It can see a page lacks a description. It cannot know the page is being retired next month |
| Explaining why you are not cited | Unreliable | Source preference differs by engine and is not visible from your codebase at all |
| Measuring whether a change worked | Unreliable | Nothing in this workflow observes outcomes. That needs instrumentation and time |
- The reliable tasks share one property: the output can be checked against the codebase. That is the whole selection rule.
- The dry-run-first, human-approves-second workflow is what makes the middle category usable. Every serious failure mode passes a build cleanly.
- Nothing here is measured against outcomes. This guide is about doing the work faster, not evidence that the work produces citations.
- The binding constraint on a team is review quality, not model capability.
- Cost figures in circulation, including the $5-15 per session estimate below, are borrowed from existing coverage. We have not published a first-party benchmark.
The safe/unsafe delegation table
Autonomy is a spectrum, not a switch. Here is where to set it, and what each row is protecting against. Concrete examples of what happens when these boundaries slip are catalogued in Claude Code SEO failure modes, which doubles as a review checklist for exactly this table.
| Capability | Sensible default | Why |
|---|---|---|
| Reading files, running builds | Allow freely | Nothing leaves the machine. The cost of asking every time exceeds the risk. |
| Editing local files on a branch | Allow within a named scope | Reviewable as a diff, revertible in one command. |
| Generating structured data | Propose only; review every field | A fabricated date, author or statistic passes a build cleanly. This is the single failure this site treats as unacceptable. |
| Editing canonicals, robots rules, redirect maps, sitemaps | Explicit approval, every time | These fail invisibly on the rendered page and are expensive to unwind once crawled. |
| Pushing, deploying, publishing | Never automatic | The only step outside the revert-and-retry model. |
| Anything touching production data | Never automatic | Same reason, with worse consequences. |
Note what the middle rows have in common. It is not that indexing directives are hard to write — they are trivial. It is that a wrong one produces a perfectly healthy-looking page while quietly removing it from an index, and nobody notices until traffic moves weeks later. Verification for those changes has to be external: fetch the live file, check the header, watch the logs. A passing build proves nothing about them, and the GPTBot versus OAI-SearchBot split is the clearest case of a one-line change with a large invisible consequence.
What Claude Code actually is, for SEO purposes
A CLI-based agentic coding tool — documented by Anthropic, with a separate web-search tool for fetching pages — that reads a codebase, proposes changes, runs builds, and with permission applies changes directly. For SEO work that means it operates on a site's actual source files: templates, meta tags, schema, sitemaps. Not advice a human then implements by hand.
That is a difference in kind from a chat assistant, not just convenience, for any task spanning dozens or hundreds of pages. It also sits on the publisher side of the wider shift toward an agent-readable web: the agent here reads your source because your site declares nothing an agent could call. It is also different from a traditional SEO crawler, which enumerates problems deterministically and then stops. The strongest setup uses both: the crawler as the source of truth about what is broken, the agent to propose and apply the specific fix. It is the method behind the zero-to-cited log study and the llms.txt test on this site.
The dry-run-first workflow
- 01 Crawl the site Read pages, sitemap, robots.txt
- 02 Identify issues Missing schema, broken links, thin meta
- 03 Propose changes Dry-run diff, nothing applied yet
- 04 Human review Approve, reject, or edit each change
- 05 Apply + verify Build passes, spot-check in browser
The critical discipline is step 3: propose changes as a reviewable diff before anything touches the live site. Alongside it, five practices do most of the safety work.
- Run the build after every change, not just at the end of a session — an early error is cheaper than debugging a large batch.
- Spot-check visually in a real browser for anything touching layout. A passing build is a syntax result, not a visual one.
- Commit incrementally so any single bad change is easy to isolate and revert.
- Never let it invent a value. A fabricated date, author or statistic passes a build cleanly. The provenance audit exists because so much published SEO material fails this test.
- Codify the rules in a file. Anything that must hold every time — design tokens, tone, what must never be invented — belongs in an instruction file the tool reads on start. A rule stated once in conversation does not survive the session.
Claude Code for SEO isn't about generating marketing copy with a prompt. It's an execution agent operating directly on a site's codebase — audits, schema, structural fixes — with a dry-run-first, human-approves-second discipline.
Share on XPrompt patterns that work and fail
The difference between a useful session and a mess is usually the instruction, not the model.
| Weak instruction | Stronger version | Why it matters |
|---|---|---|
| "Improve the SEO on this site" | "List pages under this directory with a missing meta description" | Names the scope and the check. Produces a reviewable list, not edits. |
| "Add schema to the blog" | "For this one post, draft Article schema using only values present on the page" | Constrains the source of every field. Makes fabrication visible. |
| "Fix the internal linking" | "Propose up to three contextual links per page, and show the sentence each would sit in" | Caps the change and forces the reasoning into the output. |
| "Make the titles better" | "Rewrite these ten titles. Show old and new side by side. Change nothing yet" | A batch shown together is the only way to catch formulaic output. |
| "Clean up the redirects" | "Read the current redirect map and report any chains or loops" | Diagnosis before treatment on a file where mistakes are expensive. |
Three habits produce most of the improvement: name the scope, ask for output before edits, cap the size of the change. A fourth is less obvious and just as useful — ask the agent what it is unsure about before it proceeds. An admission of uncertainty is far cheaper to handle than a confident wrong value.
The first hour, step by step
- 01 Work on a branch Everything else assumes a bad change is one command from disappearing.
- 02 Ask for a description "Describe how meta descriptions are generated here" — tests understanding before any edit.
- 03 Check it against what you know Subtly wrong here means subtly wrong in a diff later. The cheapest reliability test there is.
- 04 One change, one file Read the diff line by line. Apply. Build. Open the page.
- 05 Write down the project rules Tokens, tone, what must never be invented. An instruction file is how rules survive the session.
- 06 Only now scale up Directory, then page type. Increase scope after each successful review pass, not before.
Step 3 is the one people skip and the one that pays. If the agent's description of your codebase is subtly wrong before it has edited anything, every diff it later proposes inherits that error. Checking a description costs a minute and tells you more about reliability than any benchmark.
A worked example: reading server logs
This is the task with the clearest advantage over a chat assistant, because a log file is large, repetitive and structured — exactly what an agent handles well.
The instruction shape. Point at the log file. Ask for a table of user-agent by request count by status code, the top requested paths per bot, and anything returning an error. Match agent strings against the published references — OpenAI's, Perplexity's, Anthropic's and Google's, or Momentic's consolidated list, cross-checked against our own bot registry.
What to check first in the output. Error responses to crawlers — a problem you can fix today, with no theory required. Then whether a bot you believe is blocked is in fact still requesting, and whether training crawlers and retrieval bots are behaving differently.
The trap. User-agent strings are self-reported and can be spoofed. A line claiming to be a major crawler may not be one; verification requires a reverse lookup on the requesting address, and an agent will not do that unless you ask.
The limit. Logs tell you what was fetched, not what was used. A crawl is not a citation, and the gap between them is the whole open question in this field — taken up in the crawl-to-citation latency study. For the shape of output worth aiming at, Paul Calvano's robots.txt and AI bots analysis is the reference, and Google's crawl-budget documentation is how to interpret request volume.
Context management on a large site
The limiting factor on real sites is not capability but how much the agent can hold in view at once.
One more constraint worth naming: do not let it optimise for a metric it cannot see. Performance is the clearest case — Core Web Vitals are field measurements with documented thresholds and a population baseline, and an agent editing code observes none of them.
What it costs, and what drives the bill
Existing coverage estimates roughly $5-15 in API credits for a session auditing hundreds of pages and proposing metadata and internal-linking fixes. That is a rough, source-dependent figure, not a controlled benchmark. We have not published a first-party cost benchmark, so treat any circulating figure as an anecdote from one site with one codebase.
What the bill responds to is more useful than a number. How much gets read, not how much gets written — an audit that reads five hundred pages to change ten is dominated by the reading, so narrowing scope is the largest single lever. Session length, because a long session re-processes accumulated context repeatedly; several short scoped sessions generally cost less than one long one. Rework, because a vague instruction that produces an unusable first pass costs twice. And file structure: large monolithic templates cost more to work with than small focused ones, for the same reason they are harder for people.
Who this workflow suits, and who it does not
Honestly, a narrower group than most coverage suggests.
- Fits: people with file-level access and version control — static sites, templated CMS setups with source access, headless stacks. Without both, the review-and-revert discipline this depends on does not exist.
- Fits: repetitive, structural work across many similar pages, where the leverage is real and review cost stays manageable.
- Fits: people who can read a diff. Every safety practice here assumes a competent reviewer.
- Fits badly: a handful of pages. Setup and review overhead can exceed the work itself.
- Does not fit: closed no-code platforms without file access, which would need a different integration entirely.
- Does not help: a site whose content is only assembled by client-side JavaScript. No amount of agent-authored markup changes what a retrieval bot can read.
Four misreadings lead somewhere bad, and they are worth naming. "So I can automate my SEO" — you can automate execution; diagnosis, prioritisation and judgement remain yours. "The dry run is optional once you trust it" — trust is not the variable, since the failures that matter are the ones nobody notices. "A passing build means the change is good" — a build is a syntax check, and every serious failure mode in the companion page passes one cleanly. "An agent-written page will get cited" — nothing here supports that, and none of the citation behaviour measured on ChatGPT or Perplexity turns on how a page was authored.
Working this way in a team
A team changes the risk profile in both directions. The good part is that review already exists: a team with mandatory code review has the single most important safeguard built into its process, and most of the discipline argued for above is simply their existing workflow.
The bad part is that reviews get lazier as volume rises. An agent produces far more change than a human author would, and a reviewer facing a forty-file diff behaves differently from one facing a four-file diff. Throughput can quietly outrun attention, and nothing in the tooling will tell you when it does.
- Label the origin of agent-assisted commits. Later investigation then narrows fast.
- Agree the autonomy boundary as a team, and write it where the tool can read it. Otherwise the risk level is set by the most permissive person.
- Keep one person accountable per session.
Our expectation is that review quality, not model capability, is the binding constraint on how safely a team can use this workflow. That is reasoning from the failure modes we have catalogued, not a measured finding.
Limitations
- Nothing here is measured against outcomes. Everything is about doing the work faster. Nothing shows the work produces citations, and a study connecting the two would change what this guide should recommend.
- No first-party cost benchmark exists. The estimate above is borrowed.
- Vertical differences are out of scope — local, commerce and YMYL sites each carry constraints this general workflow does not address.
- This is a workflow description, not a guaranteed cost or outcome for any specific site.
Take the one task off the agent that it fails silently: generate your structured data in the schema generator, where every field is one you typed, then have the agent place it. A fabricated date or author inside JSON-LD passes a build cleanly, which is why the delegation table above marks that row "propose only".
Then read the review checklist that makes the rest of the table enforceable: the failure modes catalogue lists what has actually gone wrong. The reusable-component version of this workflow is covered in open-source SEO agent tooling.
Namdev, R. (2026). Claude Code for SEO: the task list (v1). Retrieved from https://ritiknamdev.com/blog/claude-code-for-seo-guide Published under CC BY 4.0 — reuse freely with attribution.
See Claude Code SEO failure modes for what goes wrong, and open-source SEO agent tooling for the reusable component approach.