Original research · Pre-registered RCT

Does an author bio affect AI citation?

Author credentials and E-E-A-T signals are recommended constantly in GEO content. No controlled test isolates whether adding a visible author bio actually changes AI citation rate. This page registers the randomised trial that would.

Ritik Namdev Ritik Namdev ·Published September 2026 ·v0 — recruiting participant sites ·16 min read ·Last verified September 2026
The short version

Author bios and credentials are recommended throughout GEO content as a citation lever, and sit near the top of most GEO advice roundups. No controlled study we could locate tests them against AI citation rate specifically. This page registers a randomised design: 300 currently-anonymous pages, half receive a named author bio, half do not, citation rate tracked before and after. Sites are being recruited; nothing has been measured.

Research status Recruiting participants

No data has been collected yet. Nothing on this page is a result. Participant sites are being enrolled. No author signals have been manipulated and no outcome measured.

What is known
  • The tactic scoreboard grades author bios and credentials Hypothesis — widely recommended, no controlled test located.
  • Google documents E-E-A-T as a quality concept for its raters; that is a different thing from a measured citation effect.
What is not yet known
  • Whether adding author credentials changes citation rate at all.
  • Which specific signal, if any, carries the effect.
  • Whether any effect differs between YMYL and non-YMYL topics.

Research question: does a byline change citation?

The question: if you add a named author bio with visible credentials to a page that currently has none, and change nothing else, does that page get cited more often by AI search engines?

Why it is unanswered: E-E-A-T's evidentiary basis is Google's rater guidelines for human quality evaluation, not a disclosed AI-citation study. Google's own AI optimization guidance and its documentation on AI features describe content quality in general terms and name no author signal specifically. The extension from "helps human raters judge quality" to "increases AI citation" is plausible and entirely untested, which is why it is graded Hypothesis on the tactic evidence scoreboard. This study is designed to move it off that grade in either direction.

What evidence exists today?

What exists is correlational and vendor-published. Zyppy has ranked candidate citation factors, Ziptie has argued that original research wins citations and described how ChatGPT appears to choose sources, and Discovered Labs has compared citation patterns across platforms. None isolates authorship, and none discloses a corpus — the same provenance problem running through every widely-quoted figure in this field and the reason the measurement standard exists. The academic GEO paper that named the field tests content-level interventions and says nothing about authorship at all.

Where did E-E-A-T actually come from?

The lineage matters for interpreting whatever this study finds. It began as E-A-T in Google's publicly released Search Quality Rater Guidelines — a document used to train human evaluators who judge search result quality, not an automated ranking signal. It is an input to checking whether Google's ranking produces good results. A fourth "Experience" dimension was added later. Turning that into an automatable AI-citation signal is an extension the industry made by analogy, not a claim any operator has validated.

How a human-rater concept became an AI-citation tactic
  1. E-A-TSearch Quality Rater Guidelines

    Expertise, Authoritativeness, Trustworthiness — a framework for training human evaluators.

    Never an automated ranking signal. An input to checking whether ranking produced good results.

  2. + ExperienceLater revision

    A fourth dimension added to the same rater framework.

    Still a human-judgement construct, still about classic search quality.

  3. GEO extensionFrom 2023 onward

    The industry extends E-E-A-T to AI citation by analogy.

    No AI-search operator has published author signals as a retrieval criterion. The extension is an assumption.

Four channels a credential could travel through

"Does E-E-A-T help?" is too coarse to test. A credential can reach a retrieval system by at least four routes, each needing a different experiment.

ChannelWhat it isTested here?
Visible textA byline and bio a human reader can see in the rendered page.Yes — this is the isolated variable.
Structured markupPerson or author properties in machine-readable schema — generate the Person JSON-LD.No — a separate study.
Off-site corroborationThe author's presence elsewhere: profiles, publications, mentions — the brand-mentions channel, which RankScience frames as a visibility gap.No — cannot be randomised.
Domain-level reputationWhether the site as a whole is treated as a credible publisher — the pattern visible in the most-cited domains in AI search.No — held constant by design.

Most public advice blends all four into one recommendation, which is exactly why that advice is unfalsifiable: a site that adds bylines, builds author profiles elsewhere and grows its domain reputation over a year cannot say which did anything. Hypothesis Our prior is that off-site corroboration is the strongest channel, because it is the one with independent evidence behind it — but that is reasoning, not data, and it is the hardest of the four to randomise.

Method: sample, control and randomisation

RCT collection pipeline
  1. 01 Recruit 300 pages Currently with no visible author byline
  2. 02 Randomize Half get a named author bio with credentials added
  3. 03 Hold content constant No other change to the page
  4. 04 Measure citation rate Before/after, treatment vs. control
  5. 05 Publish either result Pre-committed

Sample. 300 pages across recruited participant sites, all currently published with no visible author byline, and all eligible for a real named author to stand behind them. Randomisation is stratified by topic category so treatment pages cannot skew toward topics with higher baseline citation.

Control. 150 pages left anonymous and otherwise untouched for the same window. The control arm is what makes a mid-study engine change survivable: if retrieval shifts under everyone, both arms move together and the comparison holds.

Variables. The manipulated variable is the presence of the standard bio template. Measured outcome is citation rate against a query set frozen at registration, collected per engine and never blended — the surfaces already disagree, as the AI Mode versus AI Overviews source delta shows. Crawler activity is logged alongside, so no page is measured before it has been refetched. The design is the same shape as the llms.txt test.

What counts as a qualifying bio?

Every treatment page receives the same template, to keep the treatment consistent and prevent post-hoc redefinition: a full name, a one-line stated credential or role relevant to the page's topic, and a single sentence of stated experience. No external verification links, no structured markup change, and no other edit to the body content. Deliberately excluded are any change to structured markup or to publication dates, both registered as their own experiments. Holding the template constant across all 150 treatment pages is what lets a measured difference be attributed to the presence of a bio rather than to how elaborate any individual bio was.

Hypothesis: the single registered prediction

H1: Adding a named author bio with visible credentials to a previously anonymous page causally increases its AI citation rate within one collection cycle, holding the underlying content identical. One primary outcome, named before collection; every other slice — by engine, topic or window — is labelled exploratory when reported.

What would a result look like?

Here is the measurement, with invented, hypothetical numbers only. Say a panel page on "how to choose a business insurance policy" previously published with no visible author name, sits at a 4% citation rate across the tracked query set. It is randomized into the treatment group and receives the standard bio template. No other change.

One collection cycle later, its citation rate measures 9%. Its matched control page, left anonymous, moves from 4% to 5% over the same window, within normal noise. If that pattern held consistently across the full 300-page panel it would support H1; if bio and non-bio pages moved indistinguishably, that would support the null instead. Both outcomes publish.

Either way, a mechanism is worth naming in advance. A positive effect would suggest visible expertise signals help a retrieval system separate a trustworthy source from a competing anonymous page. A null would be consistent with current-generation retrieval not parsing bio content as a distinct signal at all. That matches the live-fetch finding that systems often extract only a page's primary visible content — the same main-content bias Onely describes from the other direction in what makes content LLM-friendly. It also sits next to the open question of whether those systems execute JavaScript at all.

Author credentials are recommended constantly as an AI-citation lever. No controlled test isolates whether adding a visible author bio actually changes citation rate. We're registering the RCT that would.

Share on X

Control: confounds this design has to defeat

Randomisation handles a great deal, not everything. These threats are stated in advance so a reader can check whether we managed them.

Concurrent site changes. Panel sites keep publishing and redesigning during the window. Randomising across sites reduces this, but a large change on one site can still move the numbers.

Engine changes mid-study. Retrieval systems update without notice. The control arm is what protects the finding, which is exactly why the study has one.

Topic and vertical imbalance. Citation behaviour differs enough between local and commerce queries that an unbalanced panel would produce a topic effect wearing a bio's clothes. Stratified randomisation is the fix.

Recrawl timing. A page not recrawled since the bio was added cannot show an effect; measuring before recrawl measures latency instead. That is why the crawl-to-citation latency work matters here and why crawler activity is logged per page. For scale, Ahrefs found conventional ranking itself takes months; there is no reason to assume citation moves faster.

Query set drift. If the tracked set changes composition between cycles, rates move for reasons unrelated to any page. The set is frozen at registration.

Analysis: measurement traps in a citation RCT

Beyond confounds are the ways to run this correctly and still report it wrongly. Each of these is a rule the analysis commits to now.

Confusing citation with trafficA citation is not a visit. Referral volumes, click-through and zero-click behaviour sit between the two, and traffic noise would swamp any bio effect.
Reporting a rate without the denominatorOne citation becoming two is a 100% increase and means almost nothing. Raw counts appear beside every rate.
Testing many outcomes, reporting the bestWith enough engines, slices and windows something looks significant by chance. Primary outcome named before collection; everything else labelled exploratory.
Treating one cycle as a trendCitation is volatile week to week. One cycle showing a difference is a lead; two or three agreeing is a finding.
Dropping pages that behaved awkwardlyPages that went offline, got redirected or stopped being crawled are reported, not quietly removed. Attrition differing between arms is itself a result.
Letting the analyst know the armWherever a citation call needs human judgement, the person making it should not know whether the page had a bio. Blinding is cheap here.

Persistence is a separate question this design does not answer: a change measured once says nothing about how long it lasts, which is the subject of the citation half-life study. A citation result is also not a traffic result, so nothing here converts directly into a business case.

Why a result would not transfer between engines

Even a clean positive would describe one pipeline at one moment. Retrieval systems differ in what part of a page they read: some may extract only the main content block and discard the surrounding furniture, which is exactly where a bio usually sits, while others process the full document. That single architectural difference could produce opposite results on the same page.

They also differ in how much they lean on an underlying index versus fetching live — Seer found 87% of SearchGPT citations matching Bing's top results, while Ahrefs reports far lower overlap with Google's. An engine inheriting its candidate set from a classic index could therefore show an E-E-A-T effect arriving second-hand from classic search rather than from AI retrieval at all. An engine fetching live with its own retrieval fleet could not.

Per-engine profiles already diverge. ChatGPT, Perplexity, Claude and Gemini each favour different source types, as Discovered Labs, Leapd and Profound all report, and the concordance work exists to measure that spread. Open question Whether a bio effect found on one engine replicates on another can only be answered by measuring each separately, which is why no blended headline number will be published.

Interpretation: what each outcome would license

If the result is…It licenses sayingIt does not license saying
Positive on all enginesAdding a real byline changed measured citation rate on this panel, in this window.That E-E-A-T is a ranking factor, or that the effect size transfers to your site.
Positive on one engine onlyOne pipeline appears to read the signal; the others do not.A single blended headline number, or a generalisation to "AI search".
NullThis template, on these pages, in this window, produced no detectable change.That credentials do not matter — three other channels remain untested.
UnderpoweredThe data cannot distinguish the effect from zero; here is the interval it cannot rule out.Anything at all about the direction of the effect.

One further misreading is worth pre-empting: "a positive result means add bios everywhere." Only where a real person genuinely stands behind the page. A byline on content nobody authored is fabrication, and it fails for reasons that have nothing to do with citation rates — most sharply in YMYL categories, where authorship carries obligations independent of any engine.

The ethics of a bio experiment

An experiment that adds author credentials carries an obligation most SEO tests do not. If the credentials are not true, the study manufactures exactly the misleading signal this site exists to argue against — the standard set out in how this publication works. So the panel rule is simple: every bio names a real person who genuinely stands behind the content, with a credential they actually hold. No invented experts, no borrowed titles. A page whose site cannot supply a real author is not eligible for the treatment arm.

That constraint costs the study something — it narrows the eligible pool and slows recruitment — and it is the correct trade, because a result obtained by fabricating credentials would be worthless even with perfect statistics. A second obligation runs to readers of the panel pages. A bio added for a study is still a bio a reader will rely on, so it has to be accurate on its own terms and stay accurate after the experiment ends.

Null results we would publish

The pre-commitment is the point of registering publicly. A null here is not a failed study; it is the more likely outcome and arguably the more useful one. We would publish treated and control pages moving indistinguishably, with the confidence interval alongside so a reader can see the range of effects the data cannot rule out. We would publish a result positive on one engine and null on the others, without averaging them into a headline. We would publish a result too underpowered to conclude anything, with the power calculation shown.

And we would publish a failure to complete: insufficient recruitment, excessive attrition, or a mid-study engine change severe enough to invalidate the window. Registering a study and reporting that it could not be run is more honest than quietly dropping it. Each outcome lands in the null results registry.

Reproduction: running a smaller version yourself

You do not have to wait for a 300-page panel to learn something, though a single-site version is weaker. Scripting it is covered in the Claude Code for SEO guide.

A single-site version you can set up in an afternoon
  1. 1 List your anonymous pages Pages with no visible author. A few dozen minimum for the exercise to be worth doing.
  2. 2 Split them with an actual random number Not your judgement about which pages deserve a bio. Judgement is how a confound enters.
  3. 3 Add a real byline to one half only A real author, no other edit, and log the date each page changed.
  4. 4 Wait for recrawl before measuring anything Check server logs for the bots in the AI bot registry, and confirm your robots.txt is not blocking them.
  5. 5 Track a fixed query set across both halves Over several weeks, recording the weeks where nothing happens as carefully as the weeks where something does.

Be honest about what that delivers. With a small sample on one site, only a large effect clears the noise: a flat result tells you very little, and a striking one tells you it is worth testing properly. The bots to watch for are listed in the AI bot user-agent registry, which also covers how often sites block them by accident.

Who this applies to, and who it does not

Publishers running anonymous contentDeciding whether the editorial cost of bylines — accountability, review processes, people willing to attach their name — is worth paying. The most relevant case.
Sites where every page already has an authorThe decision is already made, for good reasons. A citation result would not change it.
Reference tables, catalogues, generated dataAuthorship is not a meaningful concept there. A human name would be decoration.
Anyone selling E-E-A-T services for AI searchA null result removes the stated justification for a widely sold recommendation — which is exactly why it is worth pre-registering rather than exploring quietly.

Schedule and raw data

Recruitment runs alongside the schema and freshness RCT panels, targeting Q1 2027 for the first collection cycle. Published with the result: the randomisation assignment list, the frozen query set, the per-page citation checks including every negative, and the crawl log confirming refetch. The bio template itself is already fixed above, so the treatment can be inspected before anyone sees an outcome. Terms used here are defined in the AI search glossary. Whether this counts as GEO, AEO or LLMO work changes nothing about the evidence.

Limitations

  • Sample: recruited volunteers, 300 pages. Participant sites self-select, which skews toward publishers already engaged with AI-search questions, and eligibility requires a real author willing to be named — narrowing the pool further, plausibly toward better-resourced editorial operations.
  • One implementation of "author bio" is tested. A named byline with a stated credential, in one fixed template. Structured Person markup, external credential verification and richer bios are untested, and credential strength is held constant rather than varied, so nothing here speaks to whether a more impressive credential changes the effect size.
  • Geography and language are uncontrolled. The panel will not be balanced across regions or languages, and author-signal conventions differ between them, so the result describes the panel's markets rather than the web's.
  • Engine coverage is limited to surfaces with a checkable citation list. Agent-facing routes — an MCP server, a WebMCP endpoint or an agentic browser — may never see a rendered bio at all and are outside the design.
  • Query selection bounds the outcome. Citation rate is measured against a query set frozen at registration; a page could gain citations for queries outside it and register as unchanged.
  • Measurement depends on recrawl, which is not controlled. Pages fetched infrequently may show no effect simply because the treatment was never seen, and attrition from that cause could differ between arms.
  • Confounders remain: freshness and off-site footprint. A bio added during any wider site refresh is impossible to separate from recency effects, and an author's off-site presence — the channel we suspect is strongest — cannot be randomised or held constant.
  • Reproducibility is bounded by panel privacy. Sites that ask not to be named appear in the raw data by identifier only, so an independent replication can check the analysis but not re-observe the same pages.
  • Generalisation is single-cycle and time-bounded. The design cannot say whether any effect decays after the first recrawl, or whether it holds once engines change, and it can say nothing about whether credential relevance to the topic matters as distinct from credential presence.
Where to go next

To act on what is already supported rather than on an untested bio effect: work through the technical GEO audit, which puts retrieval access and content structure ahead of author signals. To see how strongly every competing tactic is evidenced: read the grades on the GEO tactic evidence scoreboard.

Get the result, or join the panel

The result — including a null — and the raw assignment and citation files go out to the newsletter when the first cycle closes. If you publish anonymous pages a real author could stand behind, participant sites are still being enrolled: there is no enrolment form, so how to reach me is on the about page. The trial sits inside the AI Citation Index programme.

How to cite this
Namdev, R. (2026). Does an author bio affect AI citation? (v1). Retrieved from https://ritiknamdev.com/blog/author-eeat-ai-citations-study

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

Registered alongside the freshness RCT and schema RCT as Tier 3 research in the AI Citation Index. See current grading in the GEO tactic evidence scoreboard.

§ References

Sources

Figures attributed to third parties above have not been independently verified unless stated otherwise.

Google Search Central — AI optimization guidedevelopers.google.com/search/docs/fundamentals/ai-optimization-guide Google Search Central — AI features and your websitedevelopers.google.com/search/docs/appearance/ai-features GEO: Generative Engine Optimization — Aggarwal et al., KDD 2024arxiv.org/abs/2311.09735 arXiv — GEO paper, full PDFarxiv.org/pdf/2311.09735 Wikipedia — Generative engine optimizationen.wikipedia.org/wiki/Generative_engine_optimization Ahrefs — Schema markup and AI citationsahrefs.com/blog/schema-ai-citations Ahrefs — AI search overlap between platformsahrefs.com/blog/ai-search-overlap Ahrefs — AI Overview citations and top-10 rankingsahrefs.com/blog/ai-overview-citations-top-10 Ahrefs — How long does it take to rank in Google?ahrefs.com/blog/how-long-does-it-take-to-rank-in-google-and-how-old-are-top-ranking-pages Search Engine Journal — AI Overview citations from top-ranking pages drop sharplywww.searchenginejournal.com/google-ai-overview-citations-from-top-ranking-pages-drop-sharply/568637 Semrush — AI Overviews studywww.semrush.com/blog/semrush-ai-overviews-study Ziptie — How original research wins AI citationsziptie.dev/blog/how-original-research-wins-ai-citations Ziptie — How does ChatGPT choose its sources?ziptie.dev/blog/how-does-chatgpt-choose-its-sources Zyppy — AI citation ranking factorssignal.zyppy.com/p/ai-citation-ranking-factors RankScience — AI citations, brand mentions and the visibility gapwww.rankscience.com/blog/ai-citations-brand-mentions-visibility-gap Onely — What makes content LLM-friendlywww.onely.com/blog/llm-friendly-content Discovered Labs — AI citation patterns across platformsdiscoveredlabs.com/blog/ai-citation-patterns-how-chatgpt-claude-and-perplexity-choose-sources Discovered Labs — How each platform cites sources differentlydiscoveredlabs.com/blog/chatgpt-claude-perplexity-and-google-ai-overviews-how-each-platform-cites-sources-differently Leapd — How ChatGPT, AI Overviews and Perplexity source informationwww.leapd.ai/blog/ai-visibility/how-chatgpt-google-ai-overviews-and-perplexity-source-information-in-2026 Profound — AI platform citation patternswww.tryprofound.com/blog/ai-platform-citation-patterns Salespeak — Content freshness in AI searchsalespeak.ai/aeo-news/content-freshness-ai-search Seer Interactive — 87% of SearchGPT citations match Bing top resultswww.seerinteractive.com/insights/87-percent-of-searchgpt-citations-match-bings-top-results
FAQ

Frequently asked questions

Isn’t E-E-A-T a Google ranking concept, not an AI citation one?
Correct. E-E-A-T — Experience, Expertise, Authoritativeness, Trustworthiness — originates in Google's Search Quality Rater Guidelines for human ranking evaluation. It does not come from any AI-search company's published citation criteria. This study tests whether one specific, visible proxy for it affects AI citation, which is a genuinely separate question.
What if adding an author bio has no effect?
That would be a real, useful null result. It would mean AI systems are not picking up on this specific visible signal, which is worth knowing rather than continuing to recommend the tactic on faith. The null is also the outcome we consider more likely.
Why test a visible bio rather than structured Person markup?
They are two different plausible routes: the rendered text a retrieval system reads, or structured markup a crawler parses separately. This study isolates the visible-bio channel first because it is the more universally recommended tactic. Testing Person markup in isolation is a natural follow-up.
How many pages would this need to detect a small effect?
That depends on the baseline citation rate and the effect size worth detecting, and the power calculation publishes with the protocol rather than being asserted here. The uncomfortable general point: citation rates are low and noisy, so detecting a small effect needs far more pages than most people assume. An underpowered test that finds nothing has not shown there is nothing.
Is there a risk this study encourages fake author bios?
It is a real risk and the protocol takes it seriously. Inventing credentials is fabrication, and on a health, legal or financial page it can be harmful and legally exposed. Every panel bio must name a real person who genuinely stands behind the content. The study tests whether a genuine byline changes citation, not whether a fabricated one does.
What happens if the effect appears on one engine but not another?
That would be one of the more interesting outcomes, reported per engine rather than averaged away. A split result would suggest the signal is read by a specific pipeline rather than being a universal property of AI retrieval — and would mean any single-engine study on this topic, including a positive one, generalises badly.
Can I run a smaller version of this on my own site?
Yes, with honest expectations. A single-site version cannot randomise across independent domains, so it cannot separate a bio effect from anything else changing on your site at the same time. Treat the result as a lead worth a bigger test, not a proof.
Should I add author bios while waiting for the result?
If a real person genuinely stands behind the page, yes — but for the reason that holds today, which is that readers trust named experts and accountability improves content quality. Say "we add bylines because readers trust named experts", not "because it boosts AI citation rate". The second claim has no evidence behind it yet.
How do you avoid measuring recrawl latency instead of the treatment?
No page is measured until server logs confirm it has actually been fetched again since the bio was added. A measurement taken before recrawl is measuring latency, not the treatment, which is why crawler activity is logged alongside citation for every page in both arms.
Can I join the panel?
Participant sites are being enrolled now. There is no enrolment form — the about page explains how to get in touch. Eligible sites need pages currently published with no visible author, a real author able to stand behind them, and the ability to leave the page otherwise unchanged for a collection cycle.
Why publish the design before running it?
Because a design published afterwards can be quietly reshaped to fit whatever the data said. Naming the primary outcome, the bio template and the query set in advance is what stops a null being rewritten as a positive on a subgroup found after the fact.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.