What happened

AI tools like ChatGPT learn to answer questions by reading vast amounts of text, and the text used to teach them is called a training corpus. Taiwan’s Ministry of Digital Affairs wants AI to understand Taiwan better, so it is building a corpus that collects Taiwanese text for AI developers to teach their models with. The government calls the goal “sovereign AI”, meaning AI that understands Taiwan. As Minister Lin Yi-jing puts it, the Chinese-language data that large international models read is still mostly in simplified Chinese, and “for AI to understand Taiwan better, Taiwan’s language, culture and values have to become part of what AI learns.”

The corpus went live at the end of last year, starting with the government’s own material. On September 15 the ministry announced it was opening the collection to the private sector, and for now it is asking publishers, e-book platforms and authors to license book summaries, preview chapters and works to it. The license is unpaid. The ministry stresses that taking part is voluntary and that authors who change their minds can withdraw.

On October 1 the United Daily News carried the publishing industry’s objections. Chen Hsia-min, who chairs the Independent Publishers Association of Taiwan, said that out of respect for authors, using their content should be paid for. The chairman of the national federation of publishers’ trade associations raised a further doubt: even if an author withdraws later, the withdrawal usually doesn’t reach an AI model that has already been trained. In other words, what the AI has already learned can’t be taken back.

Why it matters to you

The two sides are arguing over what content to teach AI with, and whether to pay for it. They argue this hard because what AI learns is what it ends up saying. That is exactly what the ministry has in mind with sovereign AI.

Taiwan is still working out what to teach its own AI, but the big international models did their learning long ago. ChatGPT, Gemini and the like learned largely from huge amounts of text published openly on the web, and your website, news stories about you and what people say about you on forums may all be among what they read. Nobody asked whether you’d license it, and nobody told you which year’s version of you they read.

That raises a question few brands have thought about: is the version of you that AI learned the one you are today? Suppose a guesthouse in Yilan turned its restaurant into a café three years ago, stopped serving dinner, and updated its website long ago. Years of old travel blogs and reviews saying “they do dinner” are still out there, though, far outnumbering anything newer. Which version the AI read, the owner has no way of knowing, and when a guest asks an AI for a guesthouse in Yilan that serves dinner, it may still put this one on the list. Even if the people who wrote those travel blogs later update their posts, the version the AI has already learned may still be there.

We call this signal distortion: the you that AI describes isn’t the you of today. In the guesthouse’s case, the AI still remembers how it used to be. The other kind happens when there’s so little about you online that the AI pieces you together from scraps and ends up with someone who doesn’t quite look like you.

What to do about it

Act now. When an AI answers a customer, besides drawing on what it learned long ago, it sometimes looks things up on the web on the spot. For the part it looks up, once you fix your site, the next lookup has a chance of finding the new information. The part it learned long ago, we expect to take a long time to replace, which is all the more reason to start now.

What it learned long ago is exactly what the publishers were worried about: once something has been learned, it can’t be taken back. By the same logic, once your old information has been learned, it only gets a chance to be replaced when the AI company retrains its model some day. Our judgment is that the clearer and more plentiful the information about the current you is online, the better the chance that the AI picks up the new information when it retrains.

As for whether and when an AI changes its answer, and whether every AI does, that only becomes clear if you go back and ask again over time, across several model updates; that’s how you find out whether the new information has actually been learned.

Getting AI to understand you and describe you correctly is what’s known as GEO, generative engine optimization. Going back to ask again every so often is exactly the kind of work our managed GEO service keeps up on your behalf: our team fixes up the content on your site, gets what other sites say about you telling the same story, and regularly goes back to check how each of the AI engines describes you and whether it’s describing the you of today. All you do is review the results.

Fill in the form and we’ll get in touch to pick a day for a demo, where one of our consultants sets aside time to show you live. We’ll start by showing you your site’s GEO audit score. That’s geoweb.tw’s free website audit, which checks whether AI can make sense of your site and scores it across 12 areas. Then we’ll have a few major AI engines answer questions customers in your line of work often ask, so you can hear whether, when they mention you, they’re describing the business you are now.


Further reading