Reference · Ongoing registry

The Null-Results Registry

Every pre-registered prediction on this site that turned out — or is predicted — to be a null result, tracked in one place. Almost nobody in this field publishes negative findings. This is where they go.

Ritik Namdev Ritik Namdev ·Published September 2026 ·Grows as studies report ·10 min read
The short version

Publication bias means positive, exciting findings get published, while negative ones quietly disappear. In AI-search research, this is close to total. This registry fights that directly. Every pre-registered null prediction on this site gets tracked here. Its real outcome gets recorded, whichever way it goes.

Why publication bias is the enemy here

A field where only positive results get published gives you a distorted picture. Every tactic looks like it works, because nobody wrote up the studies that found it didn't.

The fix: register a null prediction before running a study. Commit to publishing regardless of what you find. This registry is where that commitment gets tracked, and where anyone can check it.

What counts as a null result

The definition matters more than it sounds, because a loose one lets almost any disappointing outcome get filed here and dilutes what the registry means.

A study whose registered guess was "no real effect." Once the study finishes, one of two things happens. Either that guess holds up, or the study finds a real surprise in the other direction.

Both outcomes get recorded here. This registry isn't stocked with only the safe, boring nulls that were easy to predict correctly.

Three kinds of null result

"Null result" gets used as one term for three genuinely different situations. They carry different weight, and collapsing them is how a weak study gets treated like a strong one.

A well-powered null. The study had enough sample to detect an effect of the size anyone would care about, and found nothing. This is the strong version. It supports the claim that any real effect is smaller than the threshold tested, which is genuinely useful information.

An underpowered null. The study found nothing, and also would not have found a moderate effect if one existed. This tells you almost nothing. It is frequently reported in the same language as the first kind, which is why sample size and power belong in any null-result write-up rather than in a footnote.

An inconclusive result. The data came back too noisy or too inconsistent to support any conclusion. Not the same as finding no effect. This one is the most honest to report and the least satisfying, which is presumably why it is the rarest to see published anywhere.

Every entry in this registry will state which of the three it is when it reports. A registry that published all three under one label would recreate, at smaller scale, exactly the ambiguity it exists to fix.

Why null results are structurally harder to publish

It helps to separate the mechanism from the outcome. Publication bias is usually described as a result. It is more useful understood as a set of ordinary incentives acting on people who are not doing anything wrong.

Here's the actual mechanism behind publication bias, not just the outcome. A vendor selling a visibility service has a direct reason to publish findings that make their own tactics look good. They have almost no reason to publish a well-designed study that found their own advice doesn't work.

An independent publication has the opposite setup. A track record of reporting null results, even ones that contradict this site's own past advice, is exactly the credibility this site is built on. That's why this registry exists as a standing commitment, not a one-time feature.

A worked example of the damage this causes

The abstract argument lands harder as a concrete scenario. This one is invented, and it describes a pattern that is almost certainly occurring right now across this field.

Picture ten different companies each quietly testing whether adding schema markup boosts AI citations. Nine find no effect and never publish anything. One finds a small positive effect, by chance, and writes a confident blog post about it.

A reader searching for "does schema help AI citations" now finds one loud "yes" and nine silences. That reader has no way to know nine other tests came back empty. This is exactly how a tactic with no real effect ends up looking proven, just because the failures never got written down.

The file-drawer problem, quantified where possible

The metaphor is old and precise. Studies that find nothing go in a drawer. Studies that find something get published. Anyone later surveying the literature sees only the drawer's contents that escaped, and concludes the effect is better established than it is.

In academic fields this has been measured, imperfectly, by comparing registered studies against published ones and counting the gap. That comparison is possible because registration exists. In AI-search research it is not possible, because almost nothing is registered in advance. The size of this field's file drawer is genuinely unknown.

What can be observed is suggestive. Across the tactics tracked on the evidence scoreboard, the overwhelming majority of published findings are positive. Tactics get recommended; almost none get reported as tested and found ineffective. In a field where a dozen widely-repeated tactics have never been controlled-tested at all, a near-total absence of published negative findings is not plausibly because everything works.

There is one prominent exception, and it is instructive. The one peer-reviewed study in the field included a deliberate negative control, keyword stuffing, and reported that it performed worse than doing nothing. That result survives in circulation mainly because it appeared inside a paper whose other findings were positive. A standalone paper reporting only that keyword stuffing does not work would likely have been much harder to place and much less quoted.

That asymmetry is the whole problem. The result's survival depended on being bundled with good news, not on its own usefulness. This registry exists so that null findings from this site's own studies do not need a positive result to travel alongside them.

Currently registered predictions

Three entries, all awaiting collection. Small, and stated plainly rather than padded. A registry is judged by what it eventually reports, not by how full the table looks at the start.

Predictions registered as null, awaiting or reporting outcome
StudyRegistered predictionStatus
Schema RCTAdding schema produces no detectable citation-rate changeAwaiting collection
Citation half-lifeA majority of cited URLs do not remain cited after 90 daysAwaiting collection
Page speed / AI citationPage speed produces no detectable citation-rate changeRegistered, not yet scheduled

Publication bias in AI-search research is close to total — positive findings get published, null ones quietly vanish. This registry tracks every pre-registered null prediction on this site and its eventual outcome, whichever way it goes.

Share on X

Not every pre-registered study here predicts a null result. Some, like the content-freshness and author-credential RCTs, register a directional guess instead. That's a genuine best guess that something real will change, not an assumption that nothing will.

Those studies aren't tracked in this specific registry. But their real results, confirming or disconfirming the guess, get published with the same commitment. You'll find them on their own dedicated pages instead.

The three registered predictions, in detail

The table above is compressed. Each entry deserves its reasoning, because a prediction is only meaningful if you can see why it was made.

Schema markup produces no detectable citation change. Registered against the prevailing industry recommendation, which is unusual and deliberate. The reasoning: an observational test across 1,885 pages found citations "barely moved," and separate live-fetch tests found five major AI systems reading only visible HTML while ignoring JSON-LD entirely. Two independent lines of evidence point the same way. Predicting the null is the honest read of that, and it means a positive result would be a genuine surprise rather than a confirmation.

A majority of cited URLs do not remain cited after 90 days. This one predicts against persistence, which is the more provocative direction. The reasoning is weaker than the schema case, frankly: no prior data exists at all, since nobody tracks citations forward. It rests on the observation that AI answers are non-deterministic run to run, which makes stability over months seem unlikely. If citations turn out to be durable, that is a reassuring finding for anyone investing in GEO, and registering the opposite prediction is what makes it credible when reported.

Page speed produces no detectable citation change. The weakest-mechanism entry, and registered partly for that reason. There is no obvious pathway by which load time would affect whether a retrieval system quotes a passage, particularly for cached or previously-fetched content. It appears on tactic lists anyway. Testing something with no plausible mechanism is a reasonable use of a null prediction: if it comes back positive, something in the current model of how retrieval works is wrong.

Note what these three have in common. Each predicts against something either widely recommended or intuitively appealing. A registry stocked with predictions that were obviously going to come back null would demonstrate nothing about the process.

How an entry gets added

The mechanism matters as much as the intention. A registry that depends on someone remembering to add entries will quietly stop being complete.

Automatically. It comes straight from the pre-registration of any study here whose main hypothesis predicts a null or negative result. There's no editorial filtering after the fact. The prediction gets locked in before data collection starts, exactly as written on each study's own page.

How to read a null result correctly

When an entry here reports, the write-up will follow a fixed shape. Knowing that shape in advance makes the eventual results easier to weigh, and it also serves as a template for reading anyone else's null findings.

What was predicted, and when. The prediction, published before collection started, with its date. Without this, a null result is just an observation. With it, it is a test that something could have failed.

What was measured, and on what sample. The intervention, the outcome metric, the number of units, the collection windows. This is what separates a well-powered null from an underpowered one.

What effect size the study could have detected. The most important line, and the one most commonly missing from published nulls anywhere. "No effect found" means very little without "and we could have detected an effect of at least this size."

What the result does not establish. Explicitly. A null on one implementation of a tactic does not rule out other implementations. A null on one engine does not generalise to others. Stating the boundary prevents the finding from being over-applied by people quoting it later.

What would change the conclusion. The conditions under which this result should be revisited. A null that nothing could overturn is not a finding, it is a position.

Four ways null results get misread

Null findings are unusually easy to misuse, in both directions. Four failures come up repeatedly.

Treating "no detectable effect" as "zero effect." The most common one. Every study has a detection threshold. An effect below it is invisible, not absent. A small but real lift from some tactic could sit under the threshold of every study run so far.

Generalising beyond what was tested. A null on Article schema added to pages that previously had none says nothing about Product schema, or about schema on pages that already had some. The specific treatment tested is the scope of the finding.

Using it to dismiss a tactic that has other justifications. If content freshness comes back null for AI citation, that is not an argument for letting pages go stale. Freshness has independent value for accuracy, for users, and for classic search. A null on one benefit does not remove the others.

Treating one null as settling the question. A single study is a single study, whichever direction it points. The appropriate response to one well-run null is to update meaningfully toward "probably no large effect," not to close the question. Replication matters as much for negative findings as positive ones.

Notice that two of these overstate the null and two understate it. Both directions are errors, and both happen, which is why the reporting shape above states the boundary explicitly rather than leaving readers to infer it.

Where this practice comes from

None of this is invented here. Pre-registration and null-result publication are established practice in fields that went through their own credibility problems and came out with better methods.

Clinical trial registration is the clearest precedent. Registries exist because pharmaceutical trials with unfavourable outcomes were disappearing, leaving a published literature that systematically overstated how well treatments worked. The fix was structural: register the trial and its endpoints before running it, so a missing result becomes visible as an absence.

Psychology's replication crisis produced a similar response. Widely-cited findings failed to reproduce at scale, and the diagnosis pointed at flexible analysis, selective reporting, and a publication system that rewarded novelty over reliability. Pre-registration and registered reports followed.

The pattern in both cases is the same. The problem was not fraud. It was ordinary incentives operating on honest people, producing a literature that leaned in one direction. The fix was not asking people to try harder. It was changing what had to be committed to in advance.

AI-search research has the same incentive structure and none of the infrastructure. Nothing here suggests anyone in the field is acting badly. It suggests the field is at the stage those others were at before they built the mechanisms, and that borrowing the mechanisms is cheaper than rediscovering why they exist.

What this means for trusting this site

The general problem with any publication claiming rigour is that the claim is unfalsifiable from outside. Anyone can say they publish inconvenient findings. Nobody can check it unless the inconvenient findings were promised in advance and are therefore missable when absent. That is the entire mechanism at work here, and it is why this page exists as a standing list rather than a paragraph in an about page.

A publication that only ever reports confirming findings is impossible to tell apart from marketing. This registry makes that distinction checkable. Anyone can come back later and confirm whether a registered prediction actually held. You don't have to take our word for it after the fact.

Why this should matter to you specifically

Everything above is about research practice. Here is the version that affects how you spend your week.

If you're deciding where to spend real time or budget on GEO tactics, a null result saves you from wasting effort on something that doesn't work. That's worth just as much as a positive result telling you what does.

Most advice you'll read elsewhere skips this step entirely. It only tells you what to do, never what was tried and quietly abandoned. This registry is where that missing half lives.

Spotting publication bias in someone else's content

This registry covers one site's studies. The more portable skill is recognising the pattern anywhere, since most of what you read about AI search will never carry a registry of its own. A few signals.

Every tactic in the list works. A guide recommending twelve tactics, all presented as effective, is describing a field where nothing has ever been tested and failed. That is not what real measurement produces. Real testing generates duds.

No stated uncertainty anywhere. Confident language on every claim, with no tactic flagged as unproven or contested, means either the author has evidence nobody else has, or they have not checked. The second is far more common.

The recommendations match what the author sells. Not disqualifying on its own, and worth weighting. A tool vendor whose research consistently validates the thing their tool measures has an incentive problem, whether or not it influenced the work.

Case studies with no failures. A published record of client results where every engagement improved is a selection of engagements, not a record of them. The interesting question is always what the unpublished ones did.

Findings that never get revisited. A figure published once and quoted for years, against products that change quarterly, suggests nobody re-checked. Re-measurement sometimes produces inconvenient answers, which is precisely why it gets skipped.

None of these prove anything about a specific piece of content. Together they are a reasonable filter for how much weight to put on advice from a given source, which matters more in a field this thin on verifiable evidence.

What committing to this actually costs

A commitment that costs nothing is not much of a commitment, so it is worth being clear what this one gives up.

It removes the option of quiet retreat. Once a prediction is published with a date, a study that comes back awkwardly cannot simply not be mentioned. The prediction sits here regardless of what the data does.

It will eventually contradict advice published elsewhere on this site. Several tactics described here as plausible are registered for testing. If a test comes back null, this site's own earlier framing gets corrected in public. That is the intended function and it is still a cost.

Null results are worse content, commercially. They attract fewer links, fewer shares, and less quoting than a confident positive finding. Committing to publish them means committing to publish material that performs worse by every ordinary metric.

It invites the obvious criticism. A site that publishes its own failures hands critics a list of them. The alternative, publishing only successes, avoids that and forfeits the reason anyone should believe the successes.

Those costs are the point. A commitment that only ever produced favourable outcomes would not be evidence of anything. The value of this registry is entirely in the entries that will be uncomfortable to publish.

What this registry should look like in two years

A useful way to judge whether this commitment is real: describe in advance what success looks like, so it can be checked later against what actually happened.

More reported outcomes than pending predictions. Right now every entry is awaiting collection. A registry that still consists entirely of pending entries in two years has failed, regardless of how good the intentions were. The measure is reported results, not registered ones.

At least one prediction that came back wrong. If every registered null returns null, that is a signal the predictions were too safe to be informative. Being wrong in public, and reporting it, is the outcome that most demonstrates the mechanism works.

At least one entry that contradicts advice published elsewhere on this site. The site currently describes several tactics as plausible. If the testing programme never produces a result that forces a correction somewhere else, the testing is probably not being done adversarially enough.

Effect-size thresholds reported alongside every null. Not just "no effect found" but "and we could have detected an effect of at least this size." This is the detail most easily dropped under time pressure, and its absence would indicate the reporting standard slipped.

Entries from studies that were inconvenient to run. The easy case is publishing a null on a tactic nobody cares about. The real test is a null on something the site has an interest in being true.

If none of those have happened by then, the honest conclusion is that this page was a statement of intent that did not survive contact with the work. Writing that criterion down now is the only way it stays checkable later.

Limitations

  • This registry only covers studies run on this site — it isn't a full tracker of null results across the entire field.
  • Early-stage: most entries above are still waiting on data collection. They aren't reporting confirmed outcomes yet.
How to cite this
Namdev, R. (2026). The Null-Results Registry (v1). Retrieved from https://ritiknamdev.com/blog/null-results-registry

Published under CC BY 4.0 — reuse freely with attribution.

Related work on this site

Every entry links back to its full pre-registration on the Citation Index or its dedicated study page.

FAQ

Frequently asked questions

Isn't it strange to publish a "registry" before any results exist?
No. This registry holds the pre-registered null predictions from studies already published on this site. It grows as those studies report real outcomes. Read it alongside those study pages, not as a standalone results dump.
Why does a null result deserve equal billing with a positive finding?
Because it's just as useful, if the study was well-designed. Knowing that content freshness doesn't move citation rate helps you decide where to spend effort. That's just as valuable as knowing that quotations do.
How is this different from just archiving old studies?
This tracks predictions made in advance, and their real outcome. It's a mechanism for accountability. It is not a list of everything this site has ever published.
What happens if a study registered as null actually finds a positive effect?
It gets reported exactly that way. On its own study page, and reflected here too. A surprise result, in either direction, is exactly what a pre-registration is designed to make impossible to quietly bury.
Does a null result prove a tactic does not work?
No, and this is the most common misreading. A null result means no effect was detected at the sample size and conditions tested. A small real effect can hide below that threshold. The honest phrasing is "no detectable effect at this scale," not "provably zero."
Why would anyone publish results that make their own advice look wrong?
Because the alternative is being indistinguishable from marketing. A publication that only ever confirms its own prior recommendations gives a reader no way to tell research from promotion. Publishing an inconvenient result is the cheapest available proof that the process is real.
How many entries should this registry have before it means anything?
Honestly, more than it has now. A registry with three pending entries is a stated intention. A registry with a dozen reported outcomes, some of which contradicted the prediction, is evidence. This page is currently at the first stage and says so.
Could this registry itself be gamed by only registering predictions likely to come back null?
It could, which is why the process section matters. Entries come automatically from any study whose main hypothesis predicts a null, with no editorial filtering afterward. The check on gaming is that the studies themselves are pre-registered publicly, so the full set of predictions is visible, not just the ones that landed here.
Ritik Namdev
Written by

Ritik Namdev

Growth · SEO · GEO

Growth marketer documenting a brand-new site's climb into Google and the AI engines - in public, with real numbers. Every tactic here is tested on real sites before it's published.

The Lab · Weekly

One experiment. Every week.

The field notes in your inbox - one thing I tested, the raw numbers behind it, and what it means for getting cited by AI.

Free forever. Unsubscribe anytime.