Say you sell solid-wood dining tables. You fixate on one question: when a customer asks AI “where can I buy a solid-wood dining table in Taipei,” will it recommend you? So you rework that product page — add a spec table, spell out the material and origin, rewrite the subheads into the phrasing a customer would actually use. You ask ten times in a guest window. You go from being mentioned twice to eight times. You screenshot it. You relax.

Then you skip the one thing that actually mattered: going back to ask “how do I choose a table for a small apartment,” “living room furniture recommendations,” “Nordic-style dining table” — the questions you used to get mentioned in last month. If you had asked, you’d have found that a few of them quietly slipped.

You think you’re optimizing. You’re moving inventory. This piece is about that invisible bill.

Someone Actually Measured This

Most people do GEO by aiming at one query at a time: I want to show up for X, so I go fix the page about X. That feels obvious — until a paper takes it apart and measures it.

IF-GEO (arXiv:2601.13938, tested mainly on GPT-4o-mini with cross-model verification on Gemini-2.0-Flash) points out a fact most GEO research ignores: a piece of content never serves just one query. It gets retrieved by several different questions at once. So the researchers ran a clean diagnostic: optimize a document for a single query, then separately measure the visibility change for two kinds of queries — the “target query” you aimed at, and the “non-target queries” the same content should also serve but that you didn’t aim at.

Target query (the one you aimed at) Non-target queries (the ones you didn’t)
Average visibility gain +0.277 +0.087
Share of outcomes that got worse 12.4% 30.6%

The target query’s gain is more than triple the non-targets’. But the sharper number is the bottom row: the query you aimed at gets worse about one time in eight; the ones you didn’t aim at get worse about one time in three. The paper has a name for this — gain redistribution. In plain terms: single-query optimization often isn’t growing the pie, it’s moving the pie from one side to the other, and the side it’s taken from is the side you were never looking at.

This isn’t armchair reasoning; it’s a measured distribution. The reason it goes unnoticed is simple: you only ever measure the query you aimed at, and that query’s number is guaranteed to look good.

Why It Happens: One Piece of Content, Conflicting Demands

The mechanism isn’t complicated. The queries a single piece of content has to serve simultaneously often make conflicting demands about what the page should look like.

A precise-intent query like “solid-wood dining table” loves specs, materials, dimensions, origin — pack the content hard and dense and you hit it dead-on. But “how do I choose furniture for a small apartment” wants trade-off guidance, situational explanation, a help-me-decide tone. Cramming the page with specs for the first query squeezes out exactly what the second one needed. You didn’t fail to optimize — you optimized for query A while de-optimizing for query B in the same stroke.

This isn’t one paper’s isolated finding. RAID G-SEO (arXiv:2508.11158, tested on the open model GLM-4-9B-0414) states the trade-off in its own framework in black and white: modeling a single query’s intent too precisely makes the content over-fit that one query, which lowers its ability to generalize to the others. Precision and generalization sit on opposite ends of a seesaw — push one all the way down and the other lifts off.

So the act of “fix one question until it ranks” has a cost baked in. And that cost is invisible by default, because your measurement window happens to frame only the one question you wanted to see succeed.

“Then I’ll Just Run Every Question Through Many Rounds” Doesn’t Save You Either

The first instinct is often: fine, I won’t do just one question — I’ll target every question and optimize each one over many rounds, and that patches it.

That road has a ceiling too, and it bites back once you hit it. MAGEO (arXiv:2604.19516, tested on GPT-5.2, Gemini-3 Pro, and Qwen-3 Max) observed a phenomenon it calls over-optimization fatigue: visibility scores peak around the fifth iteration, then not only stop rising but start slightly dragging down the content’s faithfulness. Their recommendation is to stop early, dynamically.

In plain terms: optimization isn’t “harder is better.” It has a point of diminishing — then negative — returns. So “grind every question into the ground” isn’t the fix. It just faithfully copies the single-point disease onto every point, with a bonus risk of content drift thrown in.

The Right Move: Look at the Whole Board First, Then Decide What to Change

Turn it around: what kind of approach dodges this bill? The direction IF-GEO points to is exactly the step single-query optimization never takes — diverge first, then converge.

Find, all at once, the several representative queries a piece of content should serve; work out the edit each one wants; then do the crucial thing — put those conflicting demands side by side and find one unified revision that’s best for the whole set of queries, not best for any single one. It even swaps out the scorecard: single-point optimization watches “how much did the target query rise,” while IF-GEO watches three whole-board numbers — how bad your worst case can get, how large your downside risk is, and on what share of queries you win or at least tie.

Here’s a practical reference on how many questions to watch: the paper found that the more queries you factor in, the better the overall result — but past roughly the fifth query, the marginal benefit starts tapering. A multi-query view doesn’t have to expand forever, but “looking at one question” and “looking at five” are two different worlds.

The verdict: single-point optimization is a zero-sum shuffle; whole-board measurement is the only real gain. The difference isn’t how hard you work, it’s whether you’re watching the whole board. Without a measurement baseline that spans multiple queries, you don’t even know which queries you’re sacrificing, let alone whether the sacrifice is worth it.

So Why Is This Hard to Do Yourself

By now the anti-DIY case isn’t “the tech is too hard” — it’s that your field of view deceives you. Three realities:

  1. You first have to know which queries a piece of content should serve, not just the one you’re aiming at right now. That itself is an inventory to take, and it usually spans many pages — it won’t surface on its own.
  2. Before you touch anything, you need a baseline that covers those queries, then you re-measure all of them afterward, so you can tell whether you grew the pie or just moved it. Watch only the query you changed and the number is always pretty — that’s the trap itself.
  3. Those queries are often scattered across different pages: the best answer to one query lives on page A, another on page B, and what you do on page A ripples into page B’s performance. This isn’t “optimize a page” anymore, it’s “coordinate a whole site.”

One person staring at one page, one question, doing it by hand, makes a specific mistake — not doing it badly, but doing it too well: polishing that one question to a shine, then losing three questions somewhere they weren’t looking. Losing sight of the rest isn’t a skill problem, it’s a field-of-view problem. Your attention can only hold one question at a time, but AI scores you by adding up your performance across queries and across pages at once. You and the engine aren’t even using the same scoreboard.

So, This Week

  1. Don’t rush to optimize any single question yet. Spend twenty minutes measuring the question you “want to be recommended for” together with three or four questions you “already show up for,” all in one guest-window pass. That’s your whole-board baseline — without it, every optimization you make afterward is a blind edit. The full method is in Screenshots Lie, Backtests Don’t.
  2. After every content change, re-measure all of those questions, not just the one you touched. Only by reading them together can you tell whether you grew the pie or shuffled it.
  3. But the full version is site-wide engineering, not something you converge by staring at one page. Enumerating every query a piece of content should serve, coordinating across pages, re-measuring multiple queries after each revision — that needs a cross-page, cross-system view. To find out which queries you’re currently sacrificing and which pages the gaps span, run a full-site analysis of your site with geoweb.tw, or send the URL to [email protected] and we’ll look and send back a concrete read.

One line to close on: it’s not that you can’t optimize — it’s that you have one pair of eyes and AI has many questions to ask. The costliest mistake here was never doing one question badly. It’s doing one question too well while nobody was watching the others.

The caveat, once: all these numbers come from controlled experiments on specific models (IF-GEO on GPT-4o-mini / Gemini-2.0-Flash, RAID G-SEO on GLM-4-9B, MAGEO on GPT-5.2, and so on). Don’t treat +0.277 or 30.6% as values that will reproduce untouched on your site. What transfers is the direction, not the decimal — the phenomenon that single-query optimization sacrifices other queries has been observed again and again across different models, and that’s the one thing to remember.


Further Reading