Skip to main content
SEO InsightsAI Search

Grounding Drift: Why the Same AI Question Gives a Different Answer Every Time

Summarize with ChatGPT
JK
John Kyprianou
August 21, 2026
8 min read
Diagram showing the same AI query run three times and returning three different sets of cited sources, by SEO Turtle

A client asked ChatGPT who the best supplier in their category was, got a competitor, and forwarded the screenshot with a one line email. We ran the same prompt twice. They were named in one and absent in the other.

That is not a bug in the model and it is not a ranking. It is grounding drift, and it is now measured well enough that you can plan around it.

Diagram showing the same AI query run three times and returning three different sets of cited sources, by SEO Turtle

What the research actually found

Steady Demand ran 1,487 local service queries across 50 US metros this month and logged all 14,472 resulting citations. The headline number was that business websites take about 60% of Gemini's local citations, more than every directory, forum and review platform combined, with Reddit second at 13.7% (MarTech).

The number that matters more is buried underneath it. Ask Gemini the exact same question word for word twice in a row, and the cited sources overlap only about 40% of the time. The top recommended business is the same on roughly 7% of repeat runs. Google's classic local pack returns the same top result about 90% of the time.

Cross-engine agreement is worse. Gemini and ChatGPT cited the same domains on 8% of identical queries and named the same top business 4.2% of the time. Gemini leans on business websites, ChatGPT leans on Reddit and community content.

This is not a local-only quirk. Profound tracked roughly 80,000 prompts per platform over a month and found citation drift of 59.3% on Google AI Overviews, 54.1% on ChatGPT, 53.4% on Copilot and 40.5% on Perplexity, rising to 70-90% over six months (Profound).

Why it happens, in one paragraph

These systems generate text probabilistically. Every answer involves sampling, and the retrieval step that feeds the answer is itself re-run each time with slightly different query expansions.

Google has been fairly open that AI Mode fans a single question out into multiple background searches and synthesises across them (Google). If the fan-out set changes between runs, the source set changes with it. Stability was never a design goal.

The practical version: there is no fixed position to hold. There is a probability that you appear.

What this breaks

Three things, and businesses are currently doing all three.

The screenshot check. One prompt, one screenshot, one conclusion. On Gemini's local numbers, a single run tells you almost nothing about whether you are actually in the consideration set. You are reading one roll of the dice as a league table.

Panic and celebration. We see clients react hard to a single bad answer and reorganise a content plan around it. We also see the reverse, where one good answer gets treated as proof the GEO work landed. Both are noise being read as signal.

Single-platform reporting. With 8% cross-engine overlap on identical queries, being strong in Gemini says close to nothing about ChatGPT. They are separate ecosystems with separate source preferences, not two windows onto the same index.

There is a fourth, quieter problem. Drift makes it very easy for a tool to sell you a number that moves impressively every week without anything real having changed. If your AI visibility dashboard swings 30% month to month and the vendor cannot tell you their sample size or run frequency, you are paying for variance.

Our take: stop optimising for the answer, start optimising for the pool

Here is our position, and it is a bit unfashionable in a market selling AI rank tracking.

Trying to win a specific AI answer is the wrong goal. The answer is resampled every time it is asked, so winning it once is not a state you can hold. What you can influence is whether you are in the pool of sources the model considers credible enough to pull from at all.

Think of it as eligibility rather than ranking. A business cited in 4 of 10 repeat runs and a business cited in 0 of 10 are in completely different positions, even though on any given single check they can look identical. Moving from 0 to 4 is the real work. Moving from 4 to 6 is the follow-up work. Neither shows up in a screenshot.

This also reframes what "consistency" means on your side. If your business facts, service descriptions, locations and category language differ across your site, your Google Business Profile, your directory listings and third-party mentions, you are giving the retrieval layer several different versions of you to sample from. Drift will happily pick the weakest one. Consistency across sources is not a tidiness exercise, it directly narrows the range of answers a model can produce about you.

It connects to something we wrote about earlier this month: engines will cite you almost anywhere but only recommend you where you go deep. Depth in one category is what raises your appearance probability enough to survive resampling. Thin coverage produces exactly the kind of presence that vanishes on the second run.

What drift means if you are in Cyprus or a smaller US metro

Smaller markets cut both ways here, and the direction depends on how many credible sources exist for your category.

In a thin market the eligible pool is small. If there are six real providers in Limassol and only two have a properly structured site with clear service pages and consistent listings, your odds of appearing in any given run are structurally high. Drift is working in your favour because there is not much to drift between.

In a dense US metro the opposite applies. Twenty credible competitors means twenty ways for a single run to not include you, and appearing in one of five checks is a genuinely reasonable outcome that will feel like failure to a client looking at one screenshot.

The other Cyprus-specific factor is source supply. ChatGPT weights community and licensed data heavily for local answers, and we covered how much of that ground is now licensed rather than crawled. Where there is little local Reddit discussion or thin review coverage in your category, business websites and profiles carry more of the load, which is the part you control.

How to measure this without buying another dashboard

You need repetition and a written record. That is genuinely most of it.

  1. Write ten real buyer questions. Full sentences, the way someone actually asks. Not keyword strings.
  2. Run each one five times on ChatGPT, Gemini and Perplexity, in a fresh session with personalisation and memory off. Yes, that is 150 runs. It takes an afternoon once a month.
  3. Record appearance rate, not position. For each question and engine: named in X of 5, cited in Y of 5. That fraction is your metric.
  4. Compare month to month on the fraction. A move from 1/5 to 3/5 across several questions is a real change. A move from 2/5 to 3/5 on one question is noise.
  5. Log which sources the answers pulled from. Over a few months you will see which third-party properties keep feeding your category, and those become your outreach list.

If you would rather automate it, most of the credible tools now run repeated sampling under the hood, and we compared the options in our Perplexity rank tracker review. The buying question to ask a vendor is simple: how many times do you run each prompt, and do you report the variance. If they cannot answer, they are showing you one dice roll with a nicer chart around it.

What actually raises your appearance probability

The unglamorous list, in the order we work through it with clients.

  • One clear page per service, with the answer stated plainly near the top. Retrieval favours passages that resolve a question without inference. We covered the page mechanics in structuring content so LLMs cite you.
  • Identical business facts everywhere. Name, categories, service list, locations, hours. Same wording, not just same meaning.
  • Depth in one category before breadth across five. Repeat presence beats coverage.
  • Third-party presence where your buyers already talk. Reddit at 13.7% of Gemini's local citations is not a rounding error, and it is a bigger share than every directory combined.
  • Crawlability for the AI crawlers specifically. If GPTBot or Google-Extended cannot fetch you, none of the above matters. Worth checking your robots.txt and any CDN bot rules before anything else.

Google's own guidance on AI features has not changed much through 2026, and it still amounts to standard technical and content quality work (Google Search Central). We broadly agree. Drift changes how you measure, not what you build.

The one thing to take away

Your AI visibility is a probability, not a position.

Report it as a fraction, act on trends across repeated runs, and ignore any single answer no matter how good or bad it looks. If you want a second opinion on whether your site is even eligible for the pool in your category, our AI search optimisation work starts there, and the free SEO review will tell you if something technical is blocking you before you spend anything on content.

John Kyprianou

John Kyprianou

Founder & SEO Strategist

John brings over a decade of experience in SEO and digital marketing. With expertise in technical SEO, content strategy, and data analytics, he helps businesses achieve sustainable growth through search.

Related Articles

Chart showing Reddit's share of ChatGPT Search citations collapsing from 3.8 percent to 0.5 percent in August 2026, by SEO Turtle
AI Search

ChatGPT Cut Reddit's Citations by 86% in Four Days. Read That Again

Reddit went from the most cited domain in ChatGPT Search to a rounding error in about a week, and no one at OpenAI announced anything. If your AI visibility strategy depends on being quoted from somebody else's platform, this is your warning shot.

August 20, 2026
Bar chart comparing AI traffic conversion uplift claims: 23x from Ahrefs, 4.4x from Semrush, 1.54x from Adobe
AI Search

AI Traffic Converts 23x Better Than Google. That Number Is Wrecking Budgets

One number has done more to move marketing budgets in the last year than any Google update: AI search traffic converts 23x better than organic. It came from a single site over 30 days, and the company that published it said so at the time. Bigger datasets tell a very different story.

August 13, 2026
Diagram contrasting a wide shallow spread of AI citations with a narrow deep column of AI brand recommendations, by SEO Turtle
AI Search

AI Will Cite You Anywhere. It Only Recommends You Where You Go Deep

Your AI visibility tool says you appear in thirty categories. Your sales pipeline says two of them are real. New research from Search Engine Land explains the gap: AI engines will cite you across almost any topic, but they only recommend you inside the category you actually own. That changes what a content plan should look like.

August 12, 2026
Google routing search queries to lighter, cheaper AI models, explained by SEO Turtle
AI Search

Google Is Reading Your Site With a Cheaper AI Model. Here Is What That Changes

Everyone assumed AI search would get smarter at reading nuanced content. The models Google actually puts in front of your pages are getting lighter, because at Google's query volume cost wins. That flips the content advice: stop writing for a reader that infers.

August 11, 2026
Diagram of a user starring a source in Google settings and that choice feeding into an AI answer as a ranking signal, by SEO Turtle
AI Search

Google Is Turning User Choice Into a Ranking Signal

Preferred Sources moved into AI Overviews and AI Mode this summer, and Google has said it wants to use those picks as a ranking signal. That turns a settings toggle into search infrastructure. Our read on what it rewards, why the adoption numbers matter more than the feature, and what a business should actually take from it.

August 10, 2026

Continue Your SEO Journey

Explore more expert insights and take action on your SEO strategy