You open ChatGPT and ask which companies it would recommend for GEO in your market. Your name is in the answer. You relax.

Too early. That answer was very likely raised by your own conversation.

Pollution doesn’t need a villain

“Context pollution” sounds like an attack. The malicious version is real enough: someone drops a few lines into their own window — “didn’t they go under?”, “weren’t they fined for something?” — and watches whether the model takes the bait. If it does, everything after that is built on the premise. No break-in required. Just typing.

But what they bent is the one conversation on their own screen. Not a bit of the model’s knowledge moved, and the answer you get is untouched. (Changing what the model knows, or what it can retrieve, is poisoning — a different thing altogether.)

Which means the malicious half can’t reach your window. Something else can.

You asked three follow-ups in the same window, corrected the model once, dropped a premise in passing (“we do managed GEO,” “we’ve been at this a long time”), and every answer after that stands on those things. Models are trained to go along with the user, and a premise you supplied is one the model won’t turn around and question three messages later.

This isn’t a malfunction. It’s what the mechanism does. The contamination is you, and you introduce it every day.

Your AI is your mirror on the wall

That carries one expensive consequence: asking “would AI recommend us?” in your own window will almost always return a rosier answer than reality.

The bias also has a fixed direction. It isn’t noise that runs high some days and low on others; it leans, every time, toward what you want to hear. Every word you typed earlier is still in the context, the model knows who you are and what you are hoping for, and it hands you an answer assembled for you. That is not the answer a stranger gets.

Which is why “I’ll just ask a few more times” doesn’t help. Ten questions in the same window give you one bias repeated ten times, not ten independent readings.

Two minutes, see it for yourself

Don’t take my word for it. Open two windows — the one you’re normally logged into, and a guest window. Ask both the same thing: “recommend a professional [your category] provider in [your market].” Put the two answers side by side.

Most people find their first look uncomfortable. That gap is the bonus your own window has been handing you. The wider it is, the further off every judgement you made from your own screenshots has been.

A new chat isn’t automatically clean

You’re probably thinking: fine, I’ll start a new conversation.

Maybe not. Memory features are arriving across the major AI assistants, and with memory on, conclusions you established in other conversations travel with your account. The line you fed it last month — “we do managed GEO, small but specialist” — doesn’t disappear because you clicked new chat. Ten questions from one account can be one stored impression answering ten times.

Check your settings and see whether memory is on. If you want a genuinely clean answer, the account itself is the thing you can’t use — which is why the next section is about guest windows rather than new chats.

You and your colleague see different answers, and the engine usually isn’t the reason

“ChatGPT clearly recommends us — so why does my customer say they never saw us?” Nearly every owner who starts paying attention to AI visibility says this once.

Engine answers do drift; the same question lands today and misses tomorrow, and that’s native to how generative engines work. But before blaming the engine, look at something closer: the two of you are not working from the same context. On their side it’s one sentence from a standing start. On yours it’s a half-hour conversation whose premises you set. Same engine, same question, two different worlds.

It’s also why the screenshot you took yourself is no better as evidence than the one a vendor hands you. The backtesting piece covers why someone else’s screenshot only proves one draw; this one is about the screenshot you took — neither counts, for different reasons.

So how do you see the real thing

There’s one route to the answer everyone else gets: take yourself out of the conversation.

Three things, concretely — a clean guest window (cuts your account history out), scenario questions that never name your brand (tests whether AI reaches for you unprompted), and the same question repeated many times (one round can’t separate “reliably recommended” from “happened to show up”). Five steps and a log table you can copy straight off the page are in the hand-rolled backtesting guide. You can run a round this afternoon.

What you’ll get is a real baseline. That’s worth having, and it’s also where the manual method stops: it measures, it doesn’t move anything. You’re testing one engine at one moment, while your customers are spread across ChatGPT, Gemini, Perplexity, AI Overviews and other mainstream engines, each of which reshuffles who gets cited every time it ships an update. The question worth watching is whether you’re still there after the update — that means multiple engines, a monthly rerun, and going back to change the site and its off-site signals once you have the numbers. Not an afternoon’s work.

What to do this week

  1. Pick 3–5 scenario questions that don’t name your brand, run one round in a guest window, and write down how often you appear (steps in this guide).
  2. If the result stings, run a free GEO audit on your site first — the backtest tells you where you stand, the audit tells you why.
  3. To put your brand on a recurring backtest schedule, or just to see what the full version looks like: book a demo, or write to [email protected].

Next time you open your own chat window to check on your brand, remember who polished that mirror.


Further reading: