# What Is GEO? The Paper That Named It Measured Citation Share, Not Recommendations

URL: https://whatmodelsays.com/what-is-geo-technical-definition-2/
Published: 2026-10-09

> At Genezio, we see GEO as getting AI assistants to recommend your brand, get your facts right and describe you as you position yourself. The paper that coined the term measured one page's share of one answer; many tools now track how often your brand is mentioned. Here's what to measure instead.

**The short version:** The way we understand GEO (generative engine optimization) at [Genezio](https://genezio.com/) is this: it's the work of getting AI assistants to recommend your brand when it fits a buyer's needs, to state your facts correctly, and to describe you the way you position yourself. GEO is usually sold with one number: "up to 40%" more visibility, from the paper that coined the term.

But that paper measured how much of one AI answer's text came from one web page, in single-question tests. Since then, a number of companies have built GEO tools that treat GEO as running prompts and measuring whether, and how often, your brand shows up. Neither a page's share of an answer nor a mention count tells you whether an assistant recommends you, gets your facts right or describes you the way you position yourself in a buyer's conversation.

There's a fair objection to start with: isn't this just SEO (search engine optimization)? [Google](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) thinks so: "From Google Search's perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO." That's half right, and the half it leaves out is what this article is about.

## What did the paper that named GEO measure? A page's share of one answer

The term comes from [Aggarwal et al., "GEO: Generative Engine Optimization"](https://arxiv.org/abs/2311.09735), first posted on arXiv in November 2023 and published at KDD 2024. The authors built a test engine. For each question, it read the top 5 Google results and wrote one answer that cited those pages.

Then they measured how much of that answer came from each page. They counted the words in the sentences citing the page, with sentences near the top weighing more. Then they scaled the results so that all cited pages' shares add up to 100%, with a sentence that cites several pages splitting its words among them. That's the page's share of the answer (call it citation share).

Here's an illustrative example, not taken from the paper. An answer has 10 sentences of similar length, and 2 of them cite only your page: your share is roughly 20%. With the position weighting, your share is higher if those are the first two sentences than if they're the last two.

To test tactics, the authors rewrote one page in different ways, such as adding quotations or statistics, and checked whether its share grew.

## What does "up to 40%" mean? A bigger share of one answer

The abstract says GEO "can boost visibility by up to 40% in generative engine responses." Visibility here means the share described above, not mentions or clicks.

Take one of the tactics tested: adding quotations to the page. Averaged over the paper's 1,000-query test set, it raised the page's share of the answer from 19.3% to 27.2%. That's a gain of 7.9 percentage points, which is about 41% in relative terms. The "40%" is that relative figure: a page went from a bit under a fifth of the answer to a bit over a quarter.

**What this means:** "up to 40%" is a bigger slice of one answer's text. It isn't 40% more recommendations, traffic or sales, and the paper doesn't claim it is. It's still a useful hint that quotable, well-sourced pages get used more.

## What don't citation share and mention counts measure? Conversations, recommendations, accuracy and perception

The authors state their scope plainly: "we focus on single-turn Generative Engines." One question, one answer, no follow-ups. They also didn't measure brand mentions, whether an answer recommended anyone, whether its facts were correct, or how it portrayed the brands it named.

Brand mentions are the gap the commercial tools moved into. [Peec AI](https://peec.ai/), for example, says it runs your prompts across [ChatGPT](https://chatgpt.com/), [Gemini](https://gemini.google.com/) and [Copilot](https://copilot.microsoft.com/) daily, and defines visibility as "how often your brand gets mentioned across AI responses." [Semrush's AI Visibility Toolkit](https://www.semrush.com/kb/1493-ai-visibility-toolkit) offers "Prompt Tracking" for "specific, high-value prompts," and [Ahrefs Brand Radar](https://ahrefs.com/brand-radar) invites you to "track your brand mentions across AI answers." Some of these tools also report position, sentiment or fact checks, but the metric their pages lead with is how often you show up.

A mention is still a different outcome from a recommendation. "Options include A, B and C" is a mention. "Given your budget and team size, I'd pick B because…" is a recommendation. Accuracy is whether what the answer says about B's price, plan limits or integrations is true.

Perception is the overall picture the answer paints of B: who it's for, what it does well and where it falls short. An assistant can mention you, cite your docs and still recommend a competitor, quote you at the wrong price, or describe you as better suited to large teams when you're built for small ones.

**The implication:** if your question is "does ChatGPT suggest us when a buyer is choosing a tool, get our facts right, and describe us the way we'd describe ourselves?", the founding GEO paper can't answer it, and a mention count alone can't either. You have to measure those outcomes directly.

## So is GEO just SEO? For getting into the answer, yes

Since the paper's metric is about being used as a source, Google's objection deserves a straight answer. Google says its generative AI features "are rooted in our core Search ranking and quality systems." Its [AI features page](https://developers.google.com/search/docs/appearance/ai-features) adds: "There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary."

So getting into the pool of pages an answer draws from is SEO work: keep doing it. What SEO metrics don't show is what the answer then says about you, and whether it picks you. That's what GEO should score.

## What should you measure instead? Recommendations, accuracy and perception in buyer-like conversations

Buyers rarely ask one question. A realistic start: "I'm a solo founder with a team of 3. Recommend a CRM that's easy to set up and under $50/month." Then come follow-ups on integrations and data export, and the shortlist from the first answer isn't necessarily the final pick.

That's why we'd make the test unit a conversation, not a single prompt. The prompt-tracking tools above describe their unit as a prompt run on a schedule. That's a reasonable proxy for a buyer's first question, but not for the follow-ups that can change the shortlist.

Here's the method we propose. Write personas (short profiles of typical buyers) with a budget, team size and needs, and run the same conversations several times, because answers vary between runs. In each run, record whether you're recommended, which facts about you are wrong, and how you're described.

That last check, perception, can be off even when every fact is right. Say an assistant gets your price, plan limits and integrations right, then sums you up as "better suited to larger sales teams" and "hard to set up." No single claim is false, but the solo founder above will move on.

So compare how models describe you (who you're for, your strengths and weaknesses, how you stack up against alternatives) with how you actually position yourself. The goal is to bring what the models say in line with that positioning.

Citations still matter here, as a diagnostic: they show which pages an answer drew on, so you know what to fix when a fact or a description is off.

Check those citations in conversations, not just single prompts, because the sources can change with the conversation. At Genezio, we simulate persona-based, multi-turn conversations, and we've observed in our experiments that the types of sources retrieved in multi-turn conversations differ from those retrieved in single-turn prompts.

## The playbook

1. **Keep your SEO foundations.** Keep key pages indexed and eligible for snippets in Google Search, so AI answers can use them as sources.
2. **Ask what a "visibility" number measures.** When a tactic or a tool reports visibility, ask: a share of what, a mention in which prompts, and measured how?
3. **Write down how you want to be described.** List who you're for, your key strengths and the facts that must be right, such as price and integrations. That's the reference you'll compare model answers against.
4. **Test buyer-like conversations repeatedly.** Use personas and realistic follow-ups, and record each run's pick, factual claims, description of you and cited sources.
5. **Use citations to fix, then retest.** Where a fact or a description is off, open the cited pages, correct what's wrong or missing on yours, and add quotations and well-sourced claims where they help. Then rerun the same conversations.

Want to see whether AI assistants recommend your brand, get your facts right and describe you the way you position yourself? At [Genezio](https://genezio.com/), we look at what models say about brands in persona-based conversations, and check the facts and sources behind those answers.
