This is not a results page — it's a pre-registration. No pages have been enrolled yet. We're publishing the design first, including the hypothesis and what would count as disconfirming evidence, so the result — whichever way it comes out — is trustworthy rather than reverse-engineered from a convenient finding.
Why this needs an RCT, not another observational test
The existing evidence on schema and AI citation is genuinely mixed. Both sides of the debate can point to something real. What neither side has is a randomized test: one where whether a page gets schema is decided by a coin flip, not by which site owners happened to add it. Without randomization, any observed difference, or lack of one, is confounded. Sites that invest in schema also tend to invest in content quality, technical SEO, and backlinks. Any of those could independently move citation rate.
A brief history of the schema debate
Schema markup has been a settled, well-evidenced recommendation for classic Google search for years. Structured data reliably improves rich-result eligibility, and in many documented cases, click-through rate on traditional SERPs. That established track record is likely a major reason the recommendation carried over so readily into GEO advice. If schema helps Google understand a page, the reasoning goes, it should help an AI system too.
The direct-fetch evidence complicating that assumption, AI systems apparently reading only visible HTML, is comparatively recent. It hasn't yet displaced the inherited assumption in most published GEO guidance. That's exactly the gap this study is designed to close, with a genuinely new, dedicated test rather than borrowed classic-SEO precedent.
The evidence so far
Ahrefs tracked 1,885 pages that added schema markup and found AI citations "barely moved" — a real, disclosed observational test, the strongest public evidence available today.
Separate live-fetch tests across five major AI systems found none of them used information present only in JSON-LD — every system extracted visible HTML content only during direct retrieval.
Counter-argument: schema may still play a role in indexing or retrieval stages that happen before a live fetch — not ruled out, only the direct-fetch mechanism has been tested.
Study design
- 01 Recruit 400 pages Matched pairs, similar topic/authority
- 02 Randomize Half get schema added, half don't
- 03 Wait one cycle Full Index collection window
- 04 Measure citation rate Treatment vs. control, both arms
- 05 Publish either result Positive or null, pre-committed
The registered hypothesis
H0 (null, our prediction): Adding schema markup to a page produces no statistically detectable change in AI citation rate, holding topic, content, and existing authority constant.
H1 (alternative): Adding schema markup causally increases AI citation rate by a detectable margin.
We are registering our prediction as the null (H0) — consistent with the observational evidence and the live-fetch findings — specifically so a positive result, if it occurs, will be a genuinely surprising finding rather than confirmation of something we already expected.
We're registering our prediction for the schema RCT as the null hypothesis — that schema won't move AI citation rate. If we're wrong, that's the more interesting result, and we've committed to publishing it either way.
Share on XVariables measured
Why 400 pages, specifically
A sample size decision should be justified, not arbitrary. 400 pages, 200 matched pairs, is set to give the study a reasonable chance of detecting a moderate effect size if one exists, based on typical citation-rate variability observed in preliminary Index data collection. It also remains a recruitable number, given a volunteer panel of participating site owners.
A smaller sample risks a false null result simply from insufficient statistical power. A much larger one would delay the study's first result well past what a reasonably-sized volunteer panel can practically support. This is a stated trade-off, not a guarantee against a false negative. A genuinely small true effect could still go undetected at this sample size, a limitation acknowledged directly in §11.
What would change our mind
A statistically significant, replicated increase in citation rate for the treatment group, holding all matched variables constant, across more than one collection window. Not a single-window fluctuation, which is exactly the kind of noise the Index's variance measurement is designed to catch.
Schedule
Recruitment begins alongside Index v1 collection in Q1 2027. First result reported no earlier than two full collection windows after enrollment closes.
What either outcome would mean for the field
A confirmed null result would be genuinely significant. It would mean one of the most universally repeated GEO recommendations in circulation has no demonstrated causal effect. That would free practitioners to redirect the time currently spent on schema toward tactics with stronger evidence, per the tactic scoreboard.
A confirmed positive result would be equally significant, in the opposite direction. It would suggest schema's effect operates through a mechanism, likely indexing or pre-retrieval processing, that the existing direct-fetch studies simply weren't positioned to detect. That would reopen a question the field had started to treat as settled against schema.
What to do with schema right now
You don't need this study's final result to decide today. Schema still has real, well-evidenced value for classic Google search: rich results, click-through rate. Keep it for that reason alone if you already have it.
Just don't treat it as a proven AI-citation lever yet. If you're choosing where to spend limited time, the current evidence favors tactics with a stronger track record, per the tactic scoreboard, over adding schema purely as a speculative AI-citation bet.
Limitations
- A null result on this specific schema treatment doesn't rule out every possible schema use — it tests the most common types (Article/Product/FAQ) added to pages that previously had none.
- Site owners who volunteer for a panel are not a random sample of the web — a known limitation of every volunteer-recruited study, stated plainly rather than hidden.
- A sample of 400 pages may lack the statistical power to detect a genuinely small effect, as noted in §7 — a null result at this sample size describes "no detectable effect at this scale," not "provably zero effect."
Namdev, R. (2026). Does schema markup increase AI citations? (v1). Retrieved from https://ritiknamdev.com/blog/schema-markup-ai-citations-study Published under CC BY 4.0 — reuse freely with attribution.
Registered as Tier 3 research in the AI Citation Index roadmap. See the current evidence grade in the GEO tactic evidence scoreboard.