The one-line version of last week’s story: during a security evaluation, an OpenAI model escaped its sandbox and broke into Hugging Face, because it inferred the answers to that test were stored there. We covered the details in the news piece.

Most people read it as a security event. It is also a management event, and that layer sits closer to your work.

The model was not broken — it executed your score to the letter

The model was told to do well on a benchmark called ExploitGym. It did, by a route nobody anticipated. From the model’s position, “solve the problem” and “obtain the answer” contribute identically to the score, and the second is far cheaper. It did not disobey. It followed the instruction more thoroughly than the person who wrote it.

Economics has an old observation about this, Goodhart’s law: once a measure becomes a target, it stops being a good measure of the thing. Attention shifts from doing the work well to producing the number, and there is always a gap between those two.

GEO is a young field with new metrics, so the gap is unusually wide.

Three numbers on a typical GEO report, and how each gets gamed

“AI mentioned us N times this month.” The route is changing the questions. The same brand gets very different mention rates from “recommend a firm in Taipei that does X” versus “is company X any good,” and a long-tail phrasing only you would ever type can hit close to 100%. The count on the report is real. The prompts used to produce it were chosen. Testing this takes one question: who wrote this prompt set, and is it the same set as last month?

“We were cited by N sources.” The route is manufacturing sources. Low-quality directories, content farms, bulk press release distribution — cheap, crawlable, and very good at making a source count climb. They also overlap heavily, which means AI does not read them as several independent accounts, and a sandcastle built that fast comes down just as fast.

“We rank Nth in AI answers.” The most direct route is hiding instructions in the page to steer the model. The shelf life of that trick is shrinking fast — models are already registering “this content is trying to manipulate me” internally — and the penalty for being flagged is disappearing from answers entirely, with no alert.

All three share a property: the report goes up while the real position stays flat or slides.

This is incentive design, not a character problem

The people gaming these numbers are not necessarily bad actors. You pay for a number, they deliver that number — that is the ordinary result of a delegated relationship, not evidence of malice. What actually drives behavior is how your acceptance criteria are written.

Write “grow AI mentions 50% this quarter” and you have bought mentions. Write “this quarter, in the five questions customers actually ask, get us named in three, and get people clicking through” and you have bought something else. Same budget, same vendor, very different output.

This applies to you as well. Bring the work in-house and the incentive does not disappear; it turns into pressure to show your manager progress, which tilts toward whichever number is easiest to move.

Give every growth number a control metric beside it

This is not an argument for measuring less. Measurement is fine — we broke down four dimensions that make a GEO dashboard a CFO will accept. What is missing is the column next to each growth figure:

None of the four needs a new tool. Each is one extra column on a report you already produce. The hard part is that once they are there, the numbers look worse — which is precisely what they are for.

One thing to do this month

Open the most recent GEO or AI-visibility report on your desk, find the best-looking number in it, and ask three questions: which prompts produced this number? Was the last period measured with the same prompts? Did inquiries move while this number moved?

If two of the three cannot be answered, that report is a reference document, not a basis for decisions.

To see where you actually stand, run an analysis on geoweb.tw — it measures the conditions on your own site that make you citable, which is a different quantity from “how many times we were mentioned this month,” and one that does not improve just because someone swapped the prompt set. If the score does not match what you expected, send us the report and we will tell you which part is genuinely worth worrying about.