What the research actually says drives AI citations
Most GEO advice is assertion. A handful of controlled studies now exist, and their findings contradict several popular recommendations. Here is what has evidence behind it.
Key takeaways
- Relevance and position in retrieved context are the dominant measured factors in citation.
- Explicit prices, recent dates and inline authoritative sources all showed measurable positive effects.
- Formatting changes alone showed weak effects, despite dominating published GEO advice.
- Retrieval manipulation is the subject of published detection benchmarks and is a short-lived tactic.
The evidence base for generative engine optimisation is small but no longer empty. The most informative single result is a 2026 factorial experiment covering 252,000 trials across six large language models and eighteen manipulated factors. Its headline finding is that relevance and position in the retrieved context dominate: moving a source higher in what the model retrieves has a larger effect than most rewrites of the source itself.
- 252,000
- trials across six models and eighteen factors in the 2026 factorial study
- 115%
- citation lift for fifth-ranked sites adding inline authoritative sources
- Weak
- measured effect of formatting changes alone on citation likelihood
What has support
- Position and relevance in retrieval. The strongest and most consistent effect. It is also the finding that ties GEO back to classic search visibility, since ranking is often how you enter the retrieved set.
- Explicit prices and recent dates. Both showed measurable effects in the controlled experiment. Concrete, checkable specifics make a passage more usable as a citation.
- Recency. Models tend to prefer the most recent version of a document matching the query over an older one, which rewards genuine updates.
- Inline citation of authoritative external sources. Associated with substantial citation gains, with the largest lift going to sites not already ranking first.
- Third-party corroboration. Expert commentary and reputable external references weigh heavily in which domains get mentioned at all.
What has weak or no support
Formatting changes on their own showed weak effects. That is worth stating plainly, because a large share of published GEO advice consists of formatting instructions: add more bullet points, use more headings, shorten paragraphs. Structure helps retrieval segment your page sensibly and it helps human readers, both of which are good reasons to do it. But reformatting a page that says nothing quotable will not make it quoted.
Similarly unsupported is the idea that adding a single file or tag produces visibility. There is no equivalent of a meta keywords tag for generative engines, and claims that one exists should be treated the way you would treat that claim about search.
What is actively risky
A parallel research line studies ranking manipulation in generative engines, specifically to characterise and detect it. Adversarial passages, hidden instructions aimed at models, and text engineered purely for extraction are the subject of published benchmarks. Anything that works by deceiving the retrieval layer is being measured by the people building defences against it, which makes it a short-lived tactic with a long-lived reputational cost.
How to read GEO advice generally
Ask three questions of any recommendation. Is there a controlled comparison behind it, or is it an assertion. Does it depend on a specific model's current behaviour, which may change next month. And would it still be a good idea if generative engines disappeared tomorrow. Advice passing the third test, clearer answers, better sourcing, honest specifics, is worth doing regardless of whether the mechanism holds.
Reformatting a page that says nothing quotable does not make it quoted.
Sources