- GEO tools promise visibility scores, prompt tracking and AI search volume estimates, but LLM answers do not behave like traditional search rankings.
- AI-referred traffic is real, but current data suggests it is still small and highly dependent on the type of page users land on.
GEO tools are trying to give marketers something familiar: rankings, visibility scores, prompt volumes and dashboards that look like SEO reports.
The problem is that AI search does not behave like Google Search.
LLM answers are not stable ranked results. A brand may be cited in one answer, ignored in the next and described differently depending on the prompt, model, session and sources used. That makes AI visibility real, but much harder to measure than traditional SEO rankings.
That is the measurement problem now facing the GEO market. Marketers want to know whether they are visible in ChatGPT, Gemini, Perplexity or Google AI Overviews. But the tools trying to answer that question often have to work without real prompt-volume data, stable ranking positions or full visibility into what users actually saw.

That tension surfaced again on 8 June 2026, when several SEO practitioners challenged whether current GEO dashboards are measuring something useful or simply turning a messy new channel into familiar-looking numbers.
The core problem: citations are not rankings
There’s a massive fucking problem with tracking AI / GEO prompt visibility.
A huge problem.
GEO prompt search volumes are bullshit volumes.
AI does not offer the stability that traditional search does because ultimately AI searches do not work the same way that traditional… pic.twitter.com/K9dEwu1wwl
— Daniel Foley Carter (@foley_seo) June 8, 2026
Daniel Foley Carter, founder of SEO-stack.io, challenged the way GEO tools present prompt visibility and AI search volume. His argument was direct: AI systems do not offer the same stability as traditional search, so prompt volumes and visibility scores can easily create a false sense of precision.
The important point is not whether brands can appear in ChatGPT, Gemini, Perplexity or Google AI Overviews. They can. The issue is whether those appearances can be measured like rankings.
In traditional SEO, a keyword has a search volume, a SERP and a rank position that can be tracked over time. In AI search, the same prompt can produce different answers depending on wording, context, user history, model updates and the sources the system decides to use.
That makes “AI visibility” a weaker and more unstable metric than a normal ranking report.
The mistake is measuring AI visibility with SEO metrics.
A citation isn’t a ranking.
A prompt isn’t a keyword.
And a click isn’t the only outcome.Many buyers now discover vendors, products, and service providers directly inside AI conversations.
The challenge isn’t tracking…— WinTraction (@WinTractionAI) June 8, 2026
WinTraction framed the problem well: a citation is not a ranking, a prompt is not a keyword and a click is not the only outcome.
That is the measurement gap GEO tools are trying to solve. Marketers want to know whether their brand appears in AI answers, how often it appears and whether competitors are being recommended instead. Those are valid questions. But the current data behind those answers is still limited.
Prompt volume is not search volume
One of the biggest weaknesses in GEO reporting is prompt volume.
Search volume in SEO is already an estimate, but it is at least based on search behavior from a platform built around queries. AI platforms do not provide the same external query-volume data to marketers.
That means GEO tools usually work from synthetic prompt sets, sampled prompts or keyword data repurposed as a proxy. Those can be useful for research, but they are not the same as real prompt demand inside ChatGPT, Gemini, Claude or Perplexity.
Prompt volume will never work the way search volume does and there’s no reliable way to check volume. Best thing we can use is SEO search volume as a loose indicator since people are starting to use LLMs the way they use Google.
— Cody Schuldt (@codylschuldt) June 8, 2026
Cody Schuldt made the more practical version of the argument: prompt volume will not work like search volume, but SEO search volume can still be used as a loose indicator because some people are starting to use LLMs the way they use Google.
That is probably the right middle ground. Prompt tracking should not be dismissed entirely. But it should be treated as directional research, not as a ranking report.
AI traffic is real, but still small
The second problem is traffic.
AI referrals are growing, but they still represent a small share of total website traffic for most sites. Ahrefs analyzed 76,000 websites and found that ChatGPT accounted for 0.21% of total website traffic, while Google organic remained close to 40%.
That does not mean AI search should be ignored. Small channels can still matter if the traffic converts well. But it does mean marketers should be careful before reorganizing an entire content strategy around AI visibility scores.
The better question is not “Are we cited in AI answers?”
The better question is: “Does AI-referred traffic actually reach the right pages and produce meaningful engagement?”
LLM traffic performs differently by page type
LLM traffic isn’t always stickier than traditional search traffic.
Sure, conversion rates are generally higher, but the way visitors interact with your website varies drastically by page type.
Check out this data from Saltbox. Tools and Demo requests get much more engagement… pic.twitter.com/bphAHIOi3g
— Neil Patel (@neilpatel) June 8, 2026
Neil Patel shared Saltbox data showing that LLM traffic does not behave the same across all page types. Tools and demo-request pages performed better with LLM visitors, while articles and product pages appeared weaker compared with traditional search traffic.
That matters because it changes the practical use case for AI visibility.
If LLM visitors are most valuable on tools, demos and decision-stage pages, then chasing citations for every blog post may be the wrong strategy. AI visibility may be more useful when it helps users reduce uncertainty and take action, not when it sends low-volume traffic to generic informational content.
This is where GEO dashboards can become misleading. A low citation score on an article may look like a problem. But if AI users were unlikely to click through to that article anyway, improving the score may not produce much business value.
What GEO tools can and cannot tell you
GEO tools are not useless. They can help brands see whether they are being mentioned in AI answers, how competitors appear and which third-party sources may influence AI-generated recommendations.
That can be useful for brand monitoring, reputation analysis and content planning.
But the limits are important:
- They do not have real platform-level prompt volume data from major AI systems.
- They cannot turn AI citations into stable rank positions.
- They often rely on synthetic prompt sets or sampled queries.
- They cannot reliably show how many real users saw a specific AI answer.
- They cannot always connect a citation to a conversion.
That means AI visibility scores should be read as signals, not as hard performance data.
A GEO score can suggest that a brand is underrepresented in AI answers. It can show how a model describes a company. It can reveal whether competitors are being cited more often. But it should not be treated like a keyword ranking report.
What marketers should measure instead
For now, the best AI-search measurement is closer to channel analysis than rank tracking.
Marketers should look at:
- AI-referred sessions in GA4.
- Which AI platforms are sending traffic.
- Which landing pages receive that traffic.
- Whether AI visitors engage, convert, request demos or use tools.
- How the brand is described across major AI systems.
- Which external sources appear to shape those descriptions.
- Whether decision-stage pages are visible in AI-generated recommendations.
This creates a more useful picture than asking whether a brand “ranks” in ChatGPT.
AI visibility matters, but the unit of measurement is not a blue-link position. It is brand presence, source influence, landing-page fit and downstream behavior.
The Query Post view
The GEO market is trying to solve a real problem: marketers need to understand how AI systems describe brands, products and sources.
But the industry is also at risk of importing old SEO language into a system where that language does not fully work.
A citation is not a ranking. A prompt is not a keyword. A visibility score is not the same as demand. And a dashboard built from synthetic prompts should not decide an entire content strategy.
The practical approach is simpler: use GEO tools for direction, not certainty.
Track whether your brand appears in AI answers. Check how competitors are being described. Understand which sources AI systems use. Measure actual AI referrals in analytics. Then compare that traffic against real outcomes, especially on tools, demos, product-led pages and high-intent landing pages.
AI search is not fake. But measuring it like traditional SEO can make the data look more precise than it really is.
That is the real story behind the GEO debate: the channel is growing, but the measurement layer is still immature.
