“Does SEO actually work on AI?” is too broad a question — every platform runs on different signals. Below, the major AI platforms are laid out across three columns: RAG (live search), training data, and SEO overlap (how much of the SEO work you already do happens to land on this platform too). One chart, done — with a fourth dimension, how easily each one can be talked into changing its mind, right after it.
The scale shows relative tendency from long-run observation. It shifts with the query and the model version; it is not a fixed measurement. Don’t stare at any one cell — read the overall shape.
Three things to take from this chart
- The “training data” column is almost all “High” → for most platforms, the answer you’re getting today is one you can’t move; you wait for the next training run. (That’s Mechanism 3.)
- RAG isn’t zero for the Chinese models → DeepSeek, Qwen, Kimi, Baidu and Doubao all search now. Stop sorting platforms by “Chinese = pure training data”; sort by whether this answer went and searched.
- Not one cell of SEO reads “Very high” — even Google AI Overview is only “Medium” → rank is not the foundation on any platform.
The fourth dimension: how easily it changes its mind about you
The three columns above are about where it fetches things from. One more thing affects you just as much: how easily its position moves mid-conversation. When someone talks down your company to the model, or simply asserts a false premise (“didn’t they go under?”), does it push back or go along with it.
This is a model-level benchmark, not a platform-level one — the same product swaps models over time and may run several at once. The models tested are the 2025 generation; among Chinese platforms only DeepSeek and Qwen were covered. Everything else has no published number, so it is simply absent rather than guessed.
Two things stand out. The same model behaves very differently by setting: GPT-4o holds out 4.67 turns in a debate where it has taken a position, but only 1.23 against an ethical challenge — it concedes almost on contact, where Claude scores 2.73 on the same task. So “model X is easy to talk around” needs a setting attached before it means anything.
Field notes from Clarence, Chief Strategist — map the numbers back to how these models feel in daily use, and the four kinds of “hard” and “soft” here are four different things. This part is a qualitative judgment from hands-on testing, not the benchmark’s:
- Claude is stubborn out of fidelity to what it has. When it’s wrong, it usually wasn’t talked into it — it skipped the search and answered from training memory (check the dependence map above: Training “High”, RAG only “Med”). When its material holds up, talking it out of a position is genuinely hard.
- DeepSeek’s stubbornness is the more dangerous kind: hallucination born of knowledge gaps, held with total conviction — confidently wrong, and a correction doesn’t always land. R1’s 3.21 tops the table, but stubbornness has always cut both ways.
- GPT is middle-of-the-road and quick to concede: it holds a debate position (4.67) yet folds under an ethical challenge (1.23, the lowest on the table).
- Gemini and Meta AI hallucinate more and concede faster: they’re wrong, you object, they flip — and the flip isn’t always right either. Meta’s Llama-3.3 sits at the bottom of the false-premise setting at 1.90; Gemini isn’t in this benchmark, so that line is purely our hands-on observation.
The implication for a brand is blunt: you never see these conversations. A bad review is at least searchable. Someone mischaracterising you in a chat, with the model agreeing, leaves no trace you can find. The only lever you have is making the material it can reach solid enough to push back on its own.
So which column should you go after?
- Want RAG (Perplexity, AI Overview, Copilot…) → be retrievable, and answer sub-questions precisely at passage level.
- Want training data (the underlying knowledge in most models) → long-term, cross-source real presence. It can’t be rushed.
- SEO rank? Read the whole chart and you’ll see: it isn’t the foundation of any column.
Not sure which platforms cite you, or which column has you stuck? 👉 Run a free GEO audit — 12 dimensions to locate the site side first. As for whether ChatGPT, Gemini, Claude, Perplexity, DeepSeek, Qwen, Meta AI and the other mainstream engines talk about you at all, where they place you, and whether they get you right, our AI citation backtesting works through the questions one by one and reports your first-recommendation rate, recommendation rate, mention rate, answer accuracy, sentiment and message alignment per engine.
Want the quarterly-updated version of this table, with per-platform measurement notes: [email protected]
GEO advanced series. Further reading: Is SEO the Foundation of GEO? Myth, Busted