What the generative engine optimization paper found

The original GEO study tested nine content-editing methods across a 10,000-query benchmark and reported visibility gains of up to 40 percent. The headline travelled much farther than the experimental conditions behind it.

A lot of later writing flattened the result into a simple claim that adding quotes, statistics, or citations makes a page 40 percent more visible in AI search. The broader generative search visibility for brands problem is more complicated because getting retrieved, getting cited, getting prominent placement, and getting traffic are different events.

Read the original generative engine optimization research paper closely, and the result is still interesting. It just answers a narrower question than many summaries suggest.

The experiment measured share inside generated answers​

The generative engine optimization study introduced GEO-bench, a collection of 10,000 queries paired with relevant web sources. For each query, the experimental setup worked with five source documents and tested how editing one source changed its presence in a generated response.

Researchers tried nine generative engine optimization techniques, including adding quotations, statistics, source citations, authoritative language, simpler wording, greater fluency, unique words, technical terms, and keyword stuffing. Quotation addition, statistics addition, and citing sources were the strongest group overall, while keyword stuffing performed poorly.

The important bit sits inside the generative engine optimization metrics. One measure counted the share of answer words attributed to a source, while a position-adjusted version gave more weight to material appearing earlier in the generated response. A separate subjective measure estimated how strongly the source appeared in the answer.

Quotation addition moved the position-adjusted score from 19.3 for the unchanged source to 27.2. The change is roughly a 41 percent relative increase, which is where the famous figure comes from. It is a lift in one visibility measure inside the experiment, not evidence of 41 percent more visitors or a 41 percent increase in organic discovery.

The original GEO benchmark study also found that results differed by subject and by the source's starting position. Some lower-ranked sources gained sharply after citations, quotations, or statistics were added, while the first-ranked source could lose share. Those generative engine optimization examples matter because the experiment was measuring competition inside a fixed answer context.

Because the source shares were normalized across the five documents, visibility was partly redistributive. One source gaining a larger slice could mean another source getting a smaller one. This makes the benchmark useful for comparing edits, but it is very different from measuring whether an AI product sends more people to a website.

The 40 percent figure needs its boundaries​

A useful generative engine optimization analysis has to separate retrieval from influence after retrieval. The benchmark started with relevant sources already supplied to the system. It did not begin by asking whether an edited page would be newly crawled, indexed, retrieved from the open web, or selected from millions of alternatives.

This distinction changes how generative engine optimization works in practice. A source can become more quotable after it enters an AI system's context without becoming more likely to enter that context in the first place. The paper demonstrated the former much more directly than the latter.

The same issue affects any generative engine optimization report that treats visibility as one number. Citation frequency, amount of attributed text, placement inside an answer, factual influence, link clicks, and conversions are separate outcomes. Improving one does not automatically improve the rest.

A July 2026 critical survey reviewed 45 studies published or released since the foundational work and drew a similar boundary. Its authors found evidence that already-retrieved content can change how it is cited or used, but no reviewed method had established a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior.

The useful finding is narrower and stronger​

The early generative engine optimization principles still hold up better when stated carefully. Specific, attributable material can affect how a generative system uses a source once the source is available to the model. Empty repetition did not show the same benefit.

Generative engine optimization history gets distorted when a controlled benchmark result turns into a universal recipe. The study did not prove that every page needs a certain number of statistics, quotations, or citations. It also did not prove that the same percentages will reproduce across modern ChatGPT, Gemini, Copilot, Perplexity, or future systems.

The strongest takeaway is less flashy but more useful. Evidence-rich editing changed source visibility under controlled conditions; the effect depended on context, and the size of the gain varied. Treat the 40 percent figure as a result from a specific benchmark, not a standing promise about AI traffic.
 

Attachments

  • What the generative engine optimization paper found.webp
    What the generative engine optimization paper found.webp
    282.3 KB · Views: 1

Sponsored

Top