Unraveling the Black Box: New Study Reveals How AI Search Agents Distribute Citations and Why Raw Metrics Lie
By Tech & SEO Insights Desk
Published September 2026
In the rapidly evolving landscape of search engine optimization (SEO), digital marketers and publishers have long obsessed over how to capture the elusive attention of artificial intelligence. As conversational search agents—powered by foundational models like OpenAI’s GPT series and xAI’s Grok—increasingly replace traditional blue links with synthesized, cited answers, understanding the mechanics of AI attribution has become a holy grail for web creators.
A newly released preprint paper, published to arXiv on September 14, offers a rare, methodical look under the hood of AI search behavior. Conducted by researchers Sriram Selvam and Anneswa Ghosh, the study dives deep into whether tweaking isolated elements of a source—such as its position in a retrieval list or its structural formatting—genuinely alters how often an AI search agent cites it, or if digital strategists are chasing phantoms.
1. Main Facts: The Anatomy of AI Citations
At its core, the study investigates whether altering a single variable of a source while keeping everything else constant impacts its likelihood of earning an AI citation. The headline takeaway is as cautionary as it is illuminating: raw statistical correlations between a page’s placement and its citation rate can be deeply misleading.
While raw numbers from the initial search transcripts showed that a top-ranked result was cited roughly twice as often as a fifth-place result, isolated testing revealed a much more complex and volatile reality. When the researchers manually swapped the positions of matched sources, the actual impact of order alone was drastically smaller—and in some follow-up tests, statistically negligible.
Furthermore, the paper explores the subtle power of formatting. Pages rewritten with structured elements—such as headings, bulleted lists, and tables—frequently captured a higher density of citation markers within an answer, functioning less like a universal "SEO trick" and more as an illustration of how language models apportion cognitive credit.
Key Parameters of the Study
- Model Tested: GPT-5.4 search agent utilizing Exa as its underlying search provider.
- Methodology: Offline replayed conversations; no live webpage edits were made during the active evaluation phases.
- Dataset: 130 common questions, yielding 113 rigorously vetted source pairs confirmed by a blinded human review process.
- Peer Review Status: Preprint (not yet peer-reviewed).
2. Chronology of the Experiment: How the Test Was Built
To isolate the variables influencing AI citations, Selvam and Ghosh designed a tightly controlled experimental framework. The chronology of their investigation reveals the painstaking steps required to interrogate a modern language model’s attribution logic.
Phase 1: Data Harvesting and Prompting
The process began by prompting the GPT-5.4 search agent to answer 130 common questions utilizing independent web searches. The researchers meticulously recorded every incoming message, query execution, and search result from the 129 successfully addressed questions.
Phase 2: Identifying Valid Source Pairs
From the resulting transcripts, the researchers filtered for pairs of pages that:
- Appeared within the same search outcomes.
- Were independently screened and confirmed as supporting the exact same underlying fact.
This screening was critical: it established a baseline where either page could be fairly cited by the model. Consequently, any credit discrepancies observed later would not be due to factual superiority, but rather to how the model distributed recognition between two equally valid sources. This filtering process left 113 genuine source pairs, which were subsequently validated by a blinded human review team (confirming 103 absolute matches).
Phase 3: The Replay and Rewriting Process
Instead of interacting with live, changing web pages, the researchers replayed each saved conversation across four distinct variations. They manipulated two main factors:
- Position: Placing one page above or below the other in the retrieval lineup.
- Formatting: Displaying the source text either as plain, unstructured paragraphs or as rewritten text featuring headings, lists, or tables.
To ensure uniformity, the alternative text versions were generated via AI rewrites. Grok 4.3 generated nearly all of the experimental variations (with GPT-5.4 utilized as a fallback for a single pair), and a separate Grok review verified that the factual integrity remained intact. Because the phrasing varied slightly between the baseline and the rewrite, the authors explicitly noted that the experiment compared structured rewrites against unstructured ones, rather than strictly isolating layout from wording.
3. Supporting Data: Raw Position Gaps vs. Controlled Swaps
The quantitative findings of the study dismantle several long-held assumptions regarding search engine placement and visibility.
The Illusion of the Raw Position Gap
In the initial, unmanipulated search calls, pages occupying the top position of the search provider’s results were cited 85.1% of the time. In contrast, pages appearing in the fifth position were cited only 42.8% of the time—a massive raw gap of 42.3 percentage points.
However, the authors emphasize that "position" in this context refers strictly to the order of the five Exa search results returned in a single query, distinct from traditional Google rankings or live web layout. Because search engines inherently bubble more relevant, authoritative pages to the top, this raw difference reflects a blend of inherent page quality and positional bias.
When the researchers experimentally forced the exact same page higher within its matched pair, the likelihood of it being cited ticked up by 7.9 percentage points. Yet, once adjusted for multiple statistical tests, this finding lost statistical significance. Furthermore, a secondary testing set consisting of 56 pairs where only the order was switched yielded an estimated position effect of 0.0 points (with a 95% confidence interval ranging from -5.4 to +5.4).
The Impact of Structured Rewrites
When evaluating formatting, pages rewritten with headings and bulleted lists received an average of 0.50 more citation markers per answer compared to their plain-paragraph counterparts (95% confidence interval: 0.20 to 0.84). Given that the test answers were densely cited—boasting a median of 29 citation markers across six documents—this represented a measurable concentration of credit.
The total volume of citations per answer did not inflate, nor did the citation count of the competing page suffer significantly. Instead, the structured text acted as a stronger anchor for the model’s attribution mechanism.
However, when measuring whether a structured layout increased the baseline probability of getting cited at all, the results showed a modest 4.5 percentage point increase (ranging from -1.4 to +10.4). The authors noted that this finding was not definitive, as their study could reliably detect effect sizes only at or above approximately 8.5 points.
The Volatility of AI Reruns
Perhaps most troubling for optimization strategists is the discovery of inherent model randomness. When the researchers re-ran 120 responses using identical inputs, the AI’s binary decision to cite or ignore the target page flipped in 15% of cases—roughly one in every seven runs. Statistical estimations suggested that approximately 45% of the variation in a single run’s outcome stems entirely from underlying model stochasticity (randomness).
4. Official Responses and Industry Context
The findings arrive amid a broader industry reckoning regarding how AI search engines determine visibility and brand recommendations.
The study’s authors summarized their overarching philosophy with a stark warning in the paper’s discussion section:
"This is an attribution-sensitivity warning, not an optimization tactic."
This sentiment echoes recent independent investigations across the digital marketing ecosystem. For instance, an industry report published by SparkToro in January revealed that major conversational platforms—including ChatGPT and Google’s AI Overviews—produced identical brand recommendations less than 1% of the time when fed the exact same prompt repeatedly.
Similarly, an Ahrefs report from May highlighted that while pages cited by AI were roughly three times more likely to incorporate JSON-LD schema markup, actively implementing schema did not reliably or predictably increase citation rates in controlled tests.
Together, these data points suggest that the digital marketing community has frequently mistaken correlation for causation, attributing AI visibility wins to specific technical tweaks when those outcomes may simply be artifacts of model variance or underlying content quality.
5. Implications for SEO, Publishers, and Future Research
As the lines between traditional search and generative AI continue to blur, the implications of Selvam and Ghosh’s preprint are profound for anyone attempting to "game" or optimize for AI citations.
1. The Death of the Single-Run Benchmark
For years, SEO professionals have relied on single-query spot-checks to determine whether their content is being surfaced by AI agents. This study proves that a single answer is a dangerously weak basis for declaring a citation "won" or "lost." With a 15% volatility rate across identical reruns, tracking tools and marketers must transition toward aggregate, multi-run testing methodologies to establish true visibility baselines.
2. Caution Regarding Optimization Dogmas
The findings caution against snake-oil optimization strategies. Simply restructuring a page into bullet points or rearranging its position in a backend retrieval queue is not a silver bullet. While structured writing may help models parse and attribute text more efficiently, it does not rewrite the fundamental requirement that the underlying content must be deemed relevant and factually sound by the retrieval engine.
3. Limitations of Offline Testing
Crucially, the authors acknowledge that their research cannot definitively confirm whether reformatting a live web page boosts citations in the wild. Because the study utilized offline replayed conversations, it excluded critical live variables such as web crawling schedules, real-time retrieval ranking algorithms, and dynamic page indexing.
Looking Ahead
To validate and expand upon these findings, the research team recommends that future studies:
- Conduct multiple test runs per scenario to account for model variance.
- Broaden the scope across a wider variety of search providers and underlying large language models.
- Monitor both granular citation counts and binary citation presence to fully understand how AI agents distribute authority.
As generative engines continue to shape the future of information discovery, studies like this serve as a vital reality check—reminding the digital ecosystem that behind every AI-generated citation lies a complex, probabilistic black box that defies simplistic optimization.
