Say you track a set of buyer prompts in ChatGPT every week, and the sites cited for your category barely move. But your buyer doesn't send one prompt. They say what they need, add a constraint, and then ask, "So which one should I get?"

Many AI visibility checks in GEO (generative engine optimization) use a single prompt as their unit. We wanted to know whether a conversation changes which sources an assistant cites. So we ran 1,855 simulated shopping conversations on ChatGPT, Claude and Perplexity, sending each buyer request two ways: as one message, and as the exact same words split into three messages. (We measured which websites were cited, not which brands were recommended or what sold.)

Short version: ask an assistant the same question as one message or as a short conversation, and it cites a different set of websites. In our tests, only about half of the 10 most-cited sites matched between the two, versus about 9 of 10 when we simply reran the same prompt. If you want to know what a buyer's conversation cites, measure with conversations.

How much does an assistant agree with itself? About 9 in 10

To call any difference real, you first need a noise floor: how much the cited sites change when nothing changes at all. For each assistant, we ran one shopping question, asked the same way, many times. Then we split those runs into two random sets and compared them.

We measured agreement as overlap: of the 10 websites cited most often in one set, the share that also appear in the other set's top 10. For ChatGPT and Claude, top-10 overlap was 90%. For Perplexity, it was 95%.

So run-to-run randomness is real, but small. If two versions of a test disagree far more than this, the difference isn't noise.

Does splitting the same words into a conversation matter? Only about 5 in 10 match

Here is one of the prompts, sent as a single message:

I live in Manhattan, New York. I want to switch my wholesale coffee bean supplier — who should I look at? I need a roaster that delivers fresh beans weekly with consistent quality. It has to work well with small independent cafés — so which roasters fit best?

In the conversation version, those three sentences became messages one, two and three. Same needs, same order. Only the split changed, and top-10 overlap fell to about half:

Assistant Rerun vs rerun One message vs conversation
ChatGPT 90% 57%
Claude 90% 50%
Perplexity 95% 53%

Rerunning a prompt kept about 9 of 10 top sources; splitting it into a conversation kept about 5.

That's a drop of 33 to 42 percentage points (the plain difference between two percentages), far more than the run-to-run noise above.

Which sources does the conversation add? Retailers and brand sites

The new sources weren't random. Amazon and Walmart each gained ground in five of our eight comparisons when the question became a conversation.

In painkillers, a pharmacy climbed while government health sites faded. Walgreens was cited in 12 of 100 single-message answers. In the conversation version, it was cited in 64 of 100. The FDA went the other way: from 60 of 100 single-message answers to 40 of 100 conversations.

Manufacturer sites showed up more too. In TVs, LG's own site appeared in 19 of 100 single-message answers. In the conversation version, it appeared in 77 of 100.

Because we recorded each turn, we could also see when sellers arrived. On Perplexity, Amazon was cited in 2 of 100 single-message painkiller answers. In the conversation, it was absent from the first and second turns, then cited in 52 of 100 third turns. Walmart on ChatGPT was absent from every first turn, then cited in 33 of 100 second turns. By the third turn, it reached 47 of 100. That's why tracking each turn matters: a test that stops at one message would conclude Amazon was barely there.

What this means: each new message adds one requirement, and the assistant answers that requirement, pulling in sources that address it. By the last message, the buyer is effectively asking "which one should I get?", so the answer turns to places that sell it.

That back-and-forth is how many people actually use AI assistants. In WildChat, a public set of about 1 million real ChatGPT conversations, about 41% went past a single exchange (Zhao et al., 2024). The average conversation ran 2.5 turns, where a turn is one message plus its reply. LMSYS-Chat-1M, a similar collection spanning 25 chat models, averages about 2 turns (Zheng et al., 2023). Those are general chats collected in 2023–24, not shopping, but they show that a short conversation is a realistic way to test, not an edge case.

What should you do about it?

  1. Measure your noise floor. Run the same prompt several times and compare the top 10 cited sites across runs.
  2. Base your GEO insights on multi-turn conversations, not single prompts. Script the 3–4 messages a buyer would actually send (need, constraint, "so which one?") and run each conversation repeatedly. Single prompts don't cover everything: some sources, like the retailers above, only show up in later turns.
  3. Record citations at every turn. Track which sites each turn cites, and compare across runs and over time. Treat changes smaller than your noise floor as randomness, and act on the ones well beyond it.
  4. Audit the last-turn sources. Check the retailer pages cited in the last turn, and the manufacturer pages that conversations add, and make sure your product facts on them are current.

How did we run the test?

We ran 1,855 simulated shopping conversations on ChatGPT, Claude and Perplexity, in three categories: a wholesale coffee roaster for a Manhattan café, an over-the-counter painkiller, and a 65-inch OLED TV. Each category was asked both ways on each assistant, except painkillers on Claude: it showed no websites for medicine questions, so there was nothing to compare. That left 8 assistant-and-category comparisons. Each was backed by roughly 60 to 190 conversations.

We measured which websites were cited, overall and per turn, and compared top-10 lists. The limits: these are simulated buyers, citations aren't recommendations or sales, and we don't have run dates to report, so treat the figures as a snapshot.

Where to see this for your own category

If your AI visibility numbers come from single prompts, they miss what buyers see later in the conversation, so measure with conversations instead. Genezio simulates repeated, persona-based multi-turn conversations around your buyer questions and shows which sources get cited at each turn.