Decoding the AI Citation Black Box: New Study Reveals Why Your Rankings and Visibility Metrics Might Be Misleading

As businesses race to optimize their web content for generative AI search engines, a new preprint study is throwing cold water on conventional optimization wisdom. Released on arXiv on September 14, researchers Sriram Selvam and Anneswa Ghosh set out to answer a burning question for digital marketers and SEO professionals: If you tweak a single element of a source—such as its layout or its position in a search engine results page (SERP)—does it fundamentally change how often an AI search agent cites it?

The findings suggest that the industry’s obsession with raw correlation metrics, search result placement, and formatting tweaks may be built on shaky foundations. While the paper has not yet undergone peer review, its deep dive into offline replayed conversations provides a sobering look at how artificial intelligence models distribute credit, highlighting the deep volatility inherent in AI search citations.


Main Facts: What the Study Found

At its core, the research investigated the mechanics of AI citations by testing a single GPT-5.4 search agent powered by Exa as its search provider. Rather than making live webpage edits—which introduce a myriad of unpredictable variables like crawling latency and changing index algorithms—the researchers relied on offline replayed conversations.

The study’s most striking revelations challenge several core assumptions held by digital marketers:

  • The Position Illusion: In raw numbers, a top-ranked result in an Exa search call was cited roughly twice as often as a result sitting in the fifth position (85.1% versus 42.8%, a massive gap of 42.3 percentage points). However, when researchers actively reversed the order of matched sources in controlled tests, the actual impact of position plummeted, measuring close to zero in follow-up trials.
  • Formatting Focus Over Total Volume: Pages rewritten with structured elements—such as clear headings and bulleted lists—did not necessarily capture a higher total volume of citations for the overall answer. Instead, they concentrated credit more heavily on the rewritten page, garnering an average of 0.50 more citation markers per answer compared to identical content written as plain paragraphs.
  • The High Volatility of AI Responses: When researchers re-ran 120 identical responses using the exact same inputs, the decision of whether or not to cite a target page flipped in 15% of the cases—roughly one out of every seven runs. The authors estimate that a staggering 45% of the variation in a single run’s citation success is driven purely by model randomness.

Chronology: How the Experiment Was Conducted

To arrive at these counterintuitive conclusions, Selvam and Ghosh designed a rigorous, multi-step testing framework to isolate variables and measure true causal relationships rather than mere correlations.

Phase 1: Prompting and Baseline Collection

The researchers began by prompting the GPT-5.4 agent to answer 130 common, everyday questions, forcing the model to perform independent web searches to formulate its responses. They successfully recorded every message and search result generated across 129 of these addressed questions, amassing a rich library of transcripts.

Phase 2: Pairing and Blinded Verification

From these transcripts, the authors searched for pairs of web pages that appeared in the same search outcomes and were independently screened as supporting the exact same underlying fact. This critical screening step ensured that either page could be fairly and accurately cited by the model.

Out of the initial pool, they identified 113 valid pairs. To maintain objectivity, a subsequent blinded human review verified 103 of these as genuine, airtight matches. Any discrepancy in how credit was ultimately apportioned between these paired sources could therefore be directly attributed to the model’s internal decision-making processes.

Phase 3: Rewriting and Replay Iterations

The researchers then replayed each saved conversation four distinct ways. They manipulated two variables:

  1. Position: Placing one page physically above or below the other within the prompt’s context.
  2. Formatting: Displaying the page’s text either as plain paragraphs, or rewriting it to include headings, lists, or structured tables.

To ensure consistency, the text versions were generated via AI rewrites—with Grok 4.3 handling the vast majority and GPT-5.4 serving as a fallback for a single pair. A separate Grok review verified that the factual integrity remained intact across versions. Because the wording varied slightly between the plain and structured versions, the authors explicitly noted that their setup compared two distinct rewrites rather than cleanly isolating isolated formatting styles.


Supporting Data: Peering Beneath the Hood

The statistical granularity of the study provides a fascinating look at the limits of AI interpretability.

When analyzing the raw position gap, the 42.3-percentage-point differential between the first and fifth positions initially suggests that ranking at the very top is paramount. However, the researchers emphasize that search providers inherently place what they deem to be the "most relevant" pages at the top. Thus, the raw difference conflates true position bias with underlying page quality.

When the researchers isolated the variable by manually moving the same page higher within its matched pair, the statistical reality shifted. The chance of being cited increased by a modest 7.9 percentage points—a finding that ultimately lacked statistical significance after accounting for multiple testing adjustments. Furthermore, in a dedicated testing subset consisting of 56 pairs where only the order was swapped, the estimated effect hovered precisely at 0.0 points (with a 95% confidence interval ranging from -5.4 to +5.4).

When examining structured text rewrites, the results were similarly nuanced. Pages formatted with headings and lists secured an average of 0.50 more citation markers per answer (with a 95% confidence interval from 0.20 to 0.84). Yet, when testing the likelihood of getting cited at all, the increase was a modest 4.5 percentage points—a result the authors openly described as inconclusive, noting their testing apparatus could only reliably detect effects of roughly 8.5 points or larger.

Interestingly, when the researchers re-ran the saved searches through Grok 4.3, the structured rewrites leaned the same way, though less than half of Grok’s initial replies adhered to the correct citation formatting standards.


Official Responses and Broader Industry Context

The findings of this preprint arrive amid a growing body of industry research questioning the validity of traditional SEO metrics when applied to generative engine optimization (GEO).

For instance, an Ahrefs report published in May revealed that web pages cited by AI models were roughly three times more likely to include JSON-LD schema markup. However, when tested practically, merely adding schema did not reliably or clearly increase AI citations. Similarly, a January report from SparkToro demonstrated that when given identical prompts repeatedly, ChatGPT and Google’s AI Overviews produced the exact same brand recommendation list less than 1% of the time.

Rather than offering a new "hack" or optimization playbook, the authors of the arXiv paper issued a vital cautionary note in their discussion section:

"This is an attribution-sensitivity warning, not an optimization tactic."

This sentiment strikes at the heart of modern digital marketing strategies. Too often, brands look at vendor reports, correlation studies, or a single isolated AI answer and conclude that a specific tactic—such as adding schema, restructuring headings, or chasing top-five placement—guarantees success.


Implications for Digital Marketers and SEO Professionals

The implications of Selvam and Ghosh’s research are profound for anyone tasked with maximizing brand visibility in AI-driven search environments.

  1. Correlation Does Not Equal Causation: Just because an AI-cited page features a specific layout, schema type, or high ranking does not mean that feature caused the citation. Without controlled tests where variables are manipulated in isolation, marketers are flying blind.
  2. The Danger of Single-Run Benchmarking: Because 15% of citation outcomes flip across identical reruns—with model randomness accounting for nearly half of the variance—relying on a single test query to evaluate GEO performance is fundamentally flawed. Marketers must run multiple iterations and measure variance before declaring a strategy successful or obsolete.
  3. Attribution Is Fragile: The study proves that AI models can be deeply fickle in how they distribute credit between equally supportive sources. Small changes in prompt execution, model updates, or random decoding choices can completely alter which brand gets the nod.

Looking Ahead

As search continues its rapid evolution from keyword-matching blue links to conversational, generative answers, the metrics used to measure success must evolve as well. The researchers recommend that future studies—and practical corporate audits—run every citation test multiple times, evaluate consistency across runs, and test a diverse array of models and search providers.

Until the underlying architectures of AI search agents become more deterministic, digital marketers would do well to heed the paper’s warning: treat optimization claims with deep skepticism, avoid chasing unverified correlations, and understand that in the world of AI search, a single answer is a remarkably weak foundation upon which to build a strategy.