The '40% visibility lift' you've heard about traces to a lab test, not real search results
If you've encountered generative engine optimization marketing materials, you've likely seen a specific number repeated widely: optimization techniques can deliver up to a 40% visibility improvement. It's a genuinely compelling figure. Tracing it back to its actual origin, and comparing it against what live campaigns actually achieve, tells a considerably more modest story.
Where the 40% figure actually comes from
The claim traces exclusively to a specific academic research paper published in November 2023, which introduced and formalized the generative engine optimization framework as a field of study. The research demonstrated that adding specific content elements, statistics, citations, authoritative quotations, could improve a page's likelihood of being referenced in generative answers by up to 40% under the study's test conditions.

Those test conditions matter enormously. The original research ran its experiments using an earlier-generation language model, evaluating synthetic, researcher-constructed test queries rather than genuine, organic search traffic from real consumers. This is a completely standard, legitimate approach for foundational academic research establishing a new field, but it's a meaningfully different environment than a live, commercial automotive search campaign operating against real consumer queries and continuously evolving production retrieval systems.
A lab result became a sales pitch. The lab conditions didn't survive the trip.
— Marqstats Analyst Team
What actually happens in live production campaigns
Controlled longitudinal tracking across live, commercial-grade conversational search engines tells a different story. Real automotive visibility gains, measured across genuine production systems handling actual consumer traffic over sustained periods, average just 14.2% - roughly a third of the widely marketed figure. This isn't a failure of the underlying optimization techniques; it reflects the genuine difference between a controlled academic test environment and the messier, more variable conditions live production search actually operates under.
Why the gap between the lab and the real world is this large
Two specific factors explain most of the difference. First, prompt volatility: real consumers ask questions in a vastly wider, less predictable variety of phrasings and specific contexts than a fixed set of researcher-constructed test queries can capture, diluting the consistency of any single optimization technique's measured effect. Second, continuous retrieval retraining: production generative search systems are regularly updated and retrained on new data, meaning the underlying retrieval and citation logic a business optimized against last month may have shifted meaningfully by the time results are actually measured, in a way a single-point academic study never has to account for.

Why the more modest, real number is still genuinely worth pursuing
It would be a mistake to read this gap as evidence that generative engine optimization doesn't work. A 14.2% average visibility gain, measured across live production systems handling genuine consumer traffic, is a real, substantial, commercially meaningful result. The issue isn't that the underlying techniques fail to deliver value - it's that marketing materials citing the 40% figure without its academic, controlled-conditions context set an unrealistic benchmark that live campaigns were never going to match, potentially making genuinely solid results look disappointing by comparison.
The counter-argument: could live results eventually catch up to the academic benchmark?
A fair question is whether the 14.2% figure represents a permanent ceiling, or whether live production results might improve over time as optimization techniques mature and generative retrieval systems stabilize, potentially closing the gap with the original academic benchmark. This is plausible, and it's reasonable to expect some improvement as the field matures and best practices become better established across the industry. What seems less likely to fully close, though, is the structural gap between a controlled test environment with fixed queries and a single model version, versus the genuinely variable, continuously evolving conditions live commercial search will always operate under. Some meaningful gap between academic benchmarks and live production results is probably a durable feature of this category, not a temporary lag that will fully disappear.
What this means for anyone evaluating GEO vendor claims
- Ask any vendor citing dramatic visibility improvement statistics to specify whether the figure comes from live production campaigns or controlled academic or simulated test conditions.
- Build internal return-on-investment expectations around the empirically observed 14.2% live-production average, not the more dramatic academic benchmark, to avoid setting unmeetable internal targets.
- Recognize that prompt volatility and continuous retrieval retraining are structural features of live generative search, not temporary implementation problems, when evaluating why measured results vary from campaign to campaign.
The full market picture
Marqstats' complete Europe automotive Generative Engine Optimization market analysis, including the full academic-to-commercial provenance tracing, is available in the linked report below.
Related reportEurope Automotive Generative Engine Optimization Market Size, Share & Forecast 2026 – 2030