This is not a results page — it's a pre-registration. No pages have been enrolled yet. We're publishing the design first, including the hypothesis and what would count as disconfirming evidence, so the result — whichever way it comes out — is trustworthy rather than reverse-engineered from a convenient finding.
No data has been collected yet. Nothing on this page is a result. Participant sites are being enrolled. No page has been randomised and no citation measured.
- Ahrefs' observational test on 1,885 pages found citations 'barely moved' after schema was added — currently the strongest evidence, and it is a null.
- Live-fetch tests have found schema is often ignored at retrieval time.
- No randomised test of schema and AI citation has ever been run.
- Whether schema has any causal effect on citation rate once confounders are removed.
- Whether any effect differs by schema type or by engine.
Why this needs an RCT, not another observational test
The existing evidence on schema and AI citation is genuinely mixed. Both sides of the debate can point to something real. What neither side has is a randomized test: one where whether a page gets schema is decided by a coin flip, not by which site owners happened to add it. Without randomization, any observed difference, or lack of one, is confounded. Sites that invest in schema also tend to invest in content quality, technical SEO, and backlinks. Any of those could independently move citation rate.
A brief history of the schema debate
Schema markup has been a settled, well-evidenced recommendation for classic Google search for years. Structured data reliably improves rich-result eligibility, and in many documented cases, click-through rate on traditional SERPs. That established track record is likely a major reason the recommendation carried over so readily into GEO advice. If schema helps Google understand a page, the reasoning goes, it should help an AI system too.
The direct-fetch evidence complicating that assumption, AI systems apparently reading only visible HTML, is comparatively recent. It hasn't yet displaced the inherited assumption in most published GEO guidance. That's exactly the gap this study is designed to close, with a genuinely new, dedicated test rather than borrowed classic-SEO precedent.
The evidence so far
Ahrefs tracked 1,885 pages that added schema markup and found AI citations "barely moved" — a real, disclosed observational test, the strongest public evidence available today.
Separate live-fetch tests across five major AI systems found none of them used information present only in JSON-LD — every system extracted visible HTML content only during direct retrieval.
Counter-argument: schema may still play a role in indexing or retrieval stages that happen before a live fetch — not ruled out, only the direct-fetch mechanism has been tested.
Study design
- 01 Recruit 400 pages Matched pairs, similar topic/authority
- 02 Randomize Half get schema added, half don't
- 03 Wait one cycle Full Index collection window
- 04 Measure citation rate Treatment vs. control, both arms
- 05 Publish either result Positive or null, pre-committed
The registered hypothesis
H0 (null, our prediction): Adding schema markup to a page produces no statistically detectable change in AI citation rate, holding topic, content, and existing authority constant.
H1 (alternative): Adding schema markup causally increases AI citation rate by a detectable margin.
We are registering our prediction as the null (H0), consistent with the observational evidence and the live-fetch findings. A positive result, if it occurs, will then be a genuinely surprising finding rather than confirmation of something we already expected.
We're registering our prediction for the schema RCT as the null hypothesis — that schema won't move AI citation rate. If we're wrong, that's the more interesting result, and we've committed to publishing it either way.
Share on XVariables measured
Why 400 pages, specifically
A sample size decision should be justified, not arbitrary. 400 pages, 200 matched pairs, gives the study a reasonable chance of detecting a moderate effect size if one exists. That figure is based on typical citation-rate variability observed in preliminary Index data collection. It also remains a recruitable number, given a volunteer panel of participating site owners.
A smaller sample risks a false null result simply from insufficient statistical power. A much larger one would delay the study's first result well past what a reasonably-sized volunteer panel can practically support. This is a stated trade-off, not a guarantee against a false negative. A genuinely small true effect could still go undetected at this sample size, a limitation acknowledged directly in the limitations below.
What would change our mind
A statistically significant, replicated increase in citation rate for the treatment group, holding all matched variables constant, across more than one collection window. Not a single-window fluctuation, which is exactly the kind of noise the Index's variance measurement is designed to catch.
Schedule
Recruitment begins alongside Index v1 collection in Q1 2027. First result reported no earlier than two full collection windows after enrollment closes.
What either outcome would mean for the field
A confirmed null result would be genuinely significant. It would mean one of the most universally repeated GEO recommendations in circulation has no demonstrated causal effect. That would free practitioners to redirect the time currently spent on schema toward tactics with stronger evidence, per the tactic scoreboard.
A confirmed positive result would be equally significant, in the opposite direction. It would suggest schema's effect operates through a mechanism, likely indexing or pre-retrieval processing, that the existing direct-fetch studies simply weren't positioned to detect. That would reopen a question the field had started to treat as settled against schema.
What to do with schema right now
Structured data remains worth maintaining for classic search rich results, which is a separate and better-evidenced benefit. Treat it as part of a technical audit rather than as a citation lever, and keep it alongside the basics that AI SEO actually rests on. Crawler-traffic surveys suggest access, not markup, is the more common failure.
You don't need this study's final result to decide today. Schema still has real, well-evidenced value for classic Google search: rich results, click-through rate. Keep it for that reason alone if you already have it.
Just don't treat it as a proven AI-citation lever yet. If you're choosing where to spend limited time, the current evidence favors tactics with a stronger track record, per the tactic scoreboard, over adding schema purely as a speculative AI-citation bet.
How schema could work, mechanically
Before testing whether structured data changes citation, it is worth stating the mechanism it would have to work through. A retrieval system must first fetch and render the page, then parse it, then decide it is worth using. Markup can only help at the parsing step, and only if the crawler reached the page at all. The crawler documentation from Google and OpenAI describes that first step; none of it describes the third.
Most arguments about schema skip the mechanism. That is the reason they never resolve. "Schema helps AI understand your page" is not a mechanism. It is a slogan. There are at least four distinct places markup could act, and they have different odds of being true.
Stage one: crawl and parse. A crawler fetches HTML. It may or may not run JavaScript. JSON-LD injected by a tag manager can be invisible to a crawler that does not render — which retrieval bots execute JavaScript at all is the prior question, and it has its own answer on this site. If your markup never survives to the parsed document, no later stage can use it. If you are writing the markup by hand, the schema generator emits server-rendered JSON-LD for the types tested here.
Stage two: index construction. A search index may store structured fields separately from body text. If it does, schema could influence which documents are retrievable for a given query. This is the stage the direct-fetch tests never touched. It is the strongest remaining argument for schema.
Stage three: retrieval and ranking. A query returns a candidate set. Schema could act here as a feature, a filter, or not at all. Nobody outside the engines can see this stage, so claims about it are unfalsifiable from the outside.
Stage four: answer composition. The model reads fetched pages and writes an answer. The live-fetch tests probed exactly this stage. They found visible HTML being used, not JSON-LD.
Schema, if it has any effect on AI citation, most likely acts at stage two rather than stage four. That is an inference from where the existing tests looked. It is not a finding.
This four-stage split matters for interpreting our own result. An RCT measures the end-to-end outcome. It cannot say which stage produced it. A positive result would tell you schema works. It would not tell you why.
What does not transfer between engines
Findings about one engine travel badly. Source mixes differ sharply between ChatGPT, Perplexity and Gemini, and cross-engine agreement is the open question that would tell us how much of any result generalises.
People treat "AI search" as one thing. It is not. A finding on one system can be worthless on another, and the reasons are structural.
Some systems fetch live pages at answer time. Others answer from a pre-built index. Some do both, depending on the query. A schema effect that lives in the index would show up in the second kind and vanish in the first.
Retrieval sources differ too. Several assistants lean on a third-party web index rather than their own crawl. In that case you are not optimising for the assistant. You are optimising for whatever index it rents.
Freshness windows differ. One system may reflect a page edit within days. Another may not reflect it for weeks. If you measure too early on the slow system, you record a null that is really a timing artifact.
Citation formats differ. Some engines link inline. Some list sources at the end. Some name a brand without linking at all. If your measurement counts only linked citations, you undercount on the systems that do not link.
This is why the study reports per-engine results as well as a pooled estimate. A pooled null across seven engines could hide a real effect on one of them. We would rather publish a messy per-engine table than a clean average that misleads.
Confounds this design removes, and the ones it does not
Randomisation is powerful, but it is not magic. It removes one specific class of problem. Being precise about which class is the difference between a credible study and a confident one.
Removed by randomisation. Selection effects. In an observational study, the sites that add schema are unusual sites. They tend to be better resourced, more technical, and more active. Those traits move citation rate on their own. A coin flip breaks that link, because the coin does not know which sites are well resourced.
Reduced by matching. Baseline differences in topic, page type and domain authority. Pairing similar pages before the flip narrows the noise the test has to see through.
Not removed: concurrent platform change. If an engine ships a retrieval update mid-window, both arms move. A difference-in-differences design absorbs a uniform shift. It does not absorb a shift that interacts with schema itself.
Not removed: participant behaviour. A site owner who knows their page is in the treatment arm may give it more attention. Blinding is difficult when the treatment is a visible code change they must make themselves.
Not removed: measurement drift. Our own prompt set and scraping method could change behaviour between windows. Versioning the collection method and publishing the diff is the only defence, and it is a partial one.
We do not currently know how to blind a self-implemented treatment at panel scale. If a reader has solved this in a comparable field study, we want to hear about it before enrollment opens.
Four common misreadings of the existing null
A null result is the easiest kind of finding to misquote. This one gets stretched in four directions, and the null results registry exists partly to keep them straight.
The Ahrefs result gets cited constantly, usually incorrectly. Four misreadings dominate.
Misreading one: "schema is useless." The result concerns AI citation only. Schema's value in classic search is a separate, better-evidenced question. Nothing in that observational test speaks to it.
Misreading two: "it proved schema does nothing." It found citations barely moved in an observational sample. That is evidence of a small or absent effect in that sample. It is not proof of zero. An observational null and a causal zero are different statements.
Misreading three: "the confounds mean we can ignore it." This is the opposite error, popular with people selling schema services. Confounding here usually biases toward finding an effect, because the sites that adopt a tactic are the sites doing everything else well. Finding no effect despite that bias makes the null more notable, not less.
Misreading four: "the live-fetch tests settled it." They tested one stage. They showed models extracting visible HTML during direct retrieval. That leaves index-stage effects entirely untested, as set out in the mechanism section above.
Good measurement means holding two things at once. The current evidence leans against schema as an AI-citation lever. The question is still genuinely open. Those are not contradictory positions.
How to run a smaller version of this yourself
You do not need a 400-page panel to learn something. You need discipline about what your smaller test can conclude. Here is a version that fits inside one site.
Step 1. Pick pages that are already being cited sometimes. A page cited zero times before and zero times after teaches you nothing. You need baseline variance to detect a change in it.
Step 2. Build matched pairs inside your own site. Pair by template, topic cluster, publication age and rough traffic. Twenty pairs is a realistic ceiling for most sites.
Step 3. Randomise with something you cannot influence. A spreadsheet random number is fine. Choosing "which page feels ready" is not randomisation.
Step 4. Write down the prompt set before you touch anything. Fix the exact queries. Fix how many times you run each. Fix what counts as a citation. Changing any of these mid-test invalidates the comparison.
Step 5. Measure a baseline over at least two rounds. One round tells you nothing about noise, and AI answers are noisy. You need to know the natural swing before you can see a signal above it.
Step 6. Apply the treatment, then wait a full re-crawl cycle. Measuring the day after you deploy measures your patience, not the engine.
Step 7. Report the difference-in-differences, not the before-and-after. Treatment change minus control change. If both arms rose equally, the platform moved, not your schema.
Step 8. Expect to conclude nothing. At twenty pairs, only a very large effect is detectable. The honest output of a small test is usually "no detectable effect at this scale." That sentence is worth writing down.
A worked example of the result table
The following numbers are invented for illustration. They are not data. No pages have been enrolled, and nothing has been measured. The point is to show the shape of the output before we have it. That way the reporting format cannot be chosen after the fact to flatter a finding.
This is the single most useful thing a control arm does. Without it, the background drift of the platform gets attributed to your change. Almost every published GEO case study lacks a control arm. That is why almost every one reports a win.
Publishing the table format in advance also constrains us. We cannot later report only the treatment arm, or only the engine that moved most. The commitment is per-engine rows plus a pooled row, with the pair count and attrition shown alongside.
Reporting a null result to a stakeholder
Suppose the study returns a null, and you have spent a quarter on structured data. Telling that story badly gets the whole measurement programme defunded. Telling it well makes the programme more credible than a win would have.
Lead with the decision, not the disappointment. "We now know where not to spend the next quarter" is a result. A test that redirects budget has paid for itself.
Say what the null does not mean. It does not mean the schema work was wasted, because the classic-search value stands. It does not mean AI visibility is unwinnable. It means one specific lever did not move one specific outcome.
Give the effect size, not just the verdict. "No detectable effect above roughly this threshold" is far more useful than "it didn't work." It tells the stakeholder what size of effect you would have caught.
Name the next test. A null is only demoralising when it ends the conversation. Pair it with the next hypothesis and the cost of testing it.
Resist the rescue narrative. There is always a way to slice the data until something looks positive. One engine, one page type, one week. That slice is how measurement programmes lose their credibility. Report it as exploratory, or do not report it.
The cost and risk of the schema bet
Schema is usually described as free. It is not free, and the real costs are worth naming.
Implementation cost is real but modest. On a templated site, one engineer can ship correct markup across a whole page type in days. On a site with hand-built pages, it is a long tail of individual edits.
Maintenance cost is the larger one. Markup drifts out of sync with the page it describes. A price changes, a review count changes, an author leaves. The structured data then quietly starts lying.
The risk of mismatched markup is not hypothetical. Structured data that contradicts visible content is a documented cause of lost rich results in classic search. So the downside case is not "no AI benefit." It is "no AI benefit, plus a search regression you caused."
Opportunity cost is the one nobody counts. The engineering week spent on markup is a week not spent on the page content itself. Content quality has a better evidence base for AI citation than markup does.
Invalid or contradictory structured data can disqualify a page from rich results in classic Google search. This is a documented behaviour of the classic search stack, independent of anything about AI citation.
None of this argues for removing schema you already have. It argues against treating schema as a cost-free default when the AI-citation case is unproven.
Open questions we cannot answer yet
A pre-registration should be honest about its own blind spots. These are the questions this design will not settle. They are listed before anyone can accuse us of avoiding them.
Does schema type matter? Pooling Article, Product and FAQ into one estimate could mask a per-type effect. We would need a far larger sample to split them.
Does schema matter more for pages with weak text? If markup substitutes for unclear prose, the effect would concentrate in badly written pages. That is testable in principle and out of scope here.
Does the effect decay? A change visible in the first window might disappear by the third. Two windows is our minimum, not a sufficient answer.
Is there an index-stage effect at all? Nobody outside the engines can observe index construction directly. An end-to-end RCT infers around it and never sees it.
We would publish a null on any of these as readily as a positive. The list exists so a later paper cannot quietly claim this study answered a question it never asked.
Limitations
- A null result on this specific schema treatment doesn't rule out every possible schema use — it tests the most common types (Article/Product/FAQ) added to pages that previously had none.
- Site owners who volunteer for a panel are not a random sample of the web — a known limitation of every volunteer-recruited study, stated plainly rather than hidden.
- A sample of 400 pages may lack the statistical power to detect a genuinely small effect, as noted in the sample-size section above. A null result at this sample size describes "no detectable effect at this scale," not "provably zero effect."
The randomisation list, the per-engine result table and the raw citation counts go out to the newsletter when the first window closes — that is how to get notified when this reports. If you run a site and would consider enrolling pages, or you have solved the blinding problem above, the contact address is on the about page. There is no enrolment form; a message is enough.
Namdev, R. (2026). Does schema markup increase AI citations? (v1). Retrieved from https://ritiknamdev.com/blog/schema-markup-ai-citations-study Published under CC BY 4.0 — reuse freely with attribution.
Registered as Tier 3 research in the AI Citation Index roadmap. See the current evidence grade in the GEO tactic evidence scoreboard.