Decoding the 15%: How "Query Fan-Out" and Russian Nesting Dolls Reveal the Future of AI-Era Search Optimization

decoding-the-15-how-query-fan-out-and-russian-nesting-dolls-reveal-the-future-of-ai-era-search-optimization

For over two decades, search engine optimization (SEO) has chased a moving target. Marketers have pivoted from keyword stuffing to semantic search, from mobile-first indexing to Core Web Vitals, and most recently, to Generative Engine Optimization (GEO). Yet, beneath the layers of modern algorithmic updates, a foundational mechanic of search behavior has remained remarkably consistent.

Long before the industry coined terms like "query fan-out" or "GEO," seasoned digital PR professionals and search strategists were quietly engineering content around a principle known colloquially as the "Russian nesting doll" approach. Today, recent empirical data from AI prompt studies and decades of search engine telemetry are validating this intuition. As AI-driven search models like ChatGPT and Google’s AI Overviews rewrite how users find information, understanding how to target the invisible, un-typed query has become the ultimate competitive advantage.


Main Facts: The Intersection of Unseen Queries and AI Fan-Out

The modern search landscape is governed by two immutable realities: the stubborn persistence of completely novel queries, and the tendency of large language models (LLMs) to automatically expand a single user prompt into a cascade of granular sub-queries.

According to Google’s historical data, 15% of all daily search queries are entirely unprecedented—phrases the search engine has never encountered before in its history. Despite the evolution of semantic algorithms and the introduction of generative AI, this metric has steadfastly refused to budge. Multiplied across billions of daily global searches, this 15% equates to hundreds of millions of daily inquiries driven by breaking news, emerging slang, newly minted product names, and unique human phrasing.

Concurrently, recent dataset research conducted by digital analyst MJ Cachón illuminates how AI models process these queries behind the scenes. When running branded prompts through ChatGPT, Cachón observed that the model automatically fires off strings of "sub-queries" that the user never actually typed. A single conversational prompt fans out into dozens of targeted, multi-word variations, progressively narrowing its focus until it matches exact, quotable phrases from publisher websites.

When lined up against Google’s revelation that AI-driven queries in the United States now run roughly three times longer than traditional searches, a clear picture emerges. Length, specificity, and semantic nesting are no longer side effects of modern search—they are the terrain itself.


Chronology: From 2003 Press Releases to the AI Search Era

To understand how search optimization arrived at query fan-out, it is helpful to trace the evolution of content strategy over the past twenty-plus years.

2003–2010s: The Era of Manual Nesting and Press Releases

Long before programmatic semantic mapping, digital PR practitioners realized that fast-moving content formats like press releases possessed an inherent SEO advantage. Because press releases were published on the exact day news broke, they captured emerging vernacular faster than traditional blog posts or static website copy.

During this era, savvy writers utilized a manual "nesting doll" technique. Rather than targeting a standalone three-word phrase (e.g., “airfare to Philadelphia”), they would anchor their copy around a four-word or five-word phrase that naturally contained the shorter term within it (e.g., “cheap airfare to Philadelphia”). By embedding the smaller phrase inside the larger one, the content became discoverable by users typing either variation. Failing to include the longer variant meant missing out entirely on users searching with greater specificity.

2019: Google Codifies the 15% Rule

In October 2019, Google introduced BERT (Bidirectional Encoder Representations from Transformers), a monumental leap forward in natural language processing. Alongside this rollout, Google officially quantified a long-standing operational reality: approximately 15% of the queries processed on any given day had never been seen before.

March 2025: The Persistence of the Unknown

At Search Central Live New York City, Google’s John Mueller revisited the 15% statistic, expressing surprise at its stubborn durability. Despite the integration of advanced LLMs capable of interpreting user intent, the fraction of unprecedented queries remained constant. As search volume grew, the absolute volume of unseen queries scaled proportionally.

August 2025: Cachón’s Query Fan-Out Dataset Study

MJ Cachón published groundbreaking research analyzing how AI systems break down user inputs. By running 189 branded prompts through ChatGPT, her study tracked 1,797 automated sub-queries. The data revealed that AI models do not search linearly; they initiate searches with broad, conversational phrasing, apply site-specific operators, and progressively hunt for exact, verifiable quotes—with quote usage multiplying 25-fold from the first sub-query to the last.

May 2026: The Rise of Extended AI Mode Queries

Google released updated usage data indicating that the average AI Mode query length in the U.S. had surged to triple the length of a legacy search query. This operational data confirmed that conversational, long-tail search behaviors were rapidly becoming the default mode of interaction for consumers.


Supporting Data: What the Metrics Tell Us

The synergy between historical SEO tactics and modern AI research is supported by several core data points:

  • The 15% Unseen Query Baseline: Consistently reported across decades, proving that human language continually outpaces pre-indexed keyword databases.
  • The 3x Length Multiplier: AI-assisted search modes encourage users to write in full, conversational paragraphs rather than fractured keywords, driving query lengths up by 300%.
  • The 25-Fold Quote Increase: Cachón’s research demonstrated that as LLMs drill down into a brand or topic, their reliance on exact-match, quoted strings increases exponentially to verify claims.
  • Average Fan-Out Breadth: A single high-level branded prompt frequently expands into roughly 10 distinct sub-queries, each targeting a different angle, attribute, or verification vector of the brand.

Official Responses and Industry Perspectives

Search engine leadership and technical SEO experts have increasingly focused their discourse on how LLMs synthesize information rather than merely matching strings.

Google’s John Mueller has repeatedly highlighted the challenge of predicting user language, noting that despite decades of algorithmic training, human unpredictability ensures that search engines will always face a massive influx of novel queries.

Meanwhile, search strategists analyzing Cachón’s findings emphasize that AI fan-out represents a mirror image of traditional long-tail optimization. While human optimizers historically built outward from seed phrases to capture variations, AI systems drill inward from conversational prompts to pinpoint verifiable facts.

In both directions—outward by human design or inward by algorithmic necessity—the underlying rule of discoverability remains unchanged: Content must contain precise linguistic variations at multiple lengths, or it risks complete invisibility.


Implications for Modern Content and SEO Strategy

For brands, publishers, and marketers navigating the age of generative search, these insights demand a fundamental shift in editorial workflows. Focusing exclusively on high-volume head terms is no longer viable. Success requires aligning with how both human users and AI models construct language.

1. Optimize for the Nested Phrase, Not Just the Seed

Content creators must look beyond basic keyword research tools to identify natural linguistic hierarchies. When targeting a core three-word or four-word term, map out the broader contextual phrases that encompass it. Structure H2 subheadings and introductory paragraphs around these longer variants. By utilizing Search Console data to uncover high-impression, low-click long-tail queries, publishers can capture traffic from variations they weren’t explicitly targeting.

2. Match the Speed of the News Cycle

Because 15% of daily queries are brand-new and closely tied to real-time events, static, quarterly content calendars are insufficient. Organizations must empower same-day publishing channels—such as rapid-response commentary, corporate newsrooms, and agile digital PR strategies—to stake a claim on brand-new terminology before competitors recognize its existence.

3. Write Literal, Quotable Answers

Because AI search engines and retrieval-augmented generation (RAG) systems verify facts by hunting for exact-match strings, content must feature clear, standalone sentences that answer specific questions. If a sentence cannot be lifted entirely out of its paragraph and still remain coherent and factually accurate, it should be rewritten.


Conclusion

The evolution from manual "Russian nesting dolls" to automated "query fan-out" proves a timeless truth in digital marketing: the fundamentals of search have always belonged to the long tail. As AI search engines continue to parse, expand, and verify the web in real time, the brands that succeed will not be those shouting the loudest head terms, but those whose content provides the precise linguistic nesting required to be found—no matter how novel the question may be.