What happened
On 2 August the transparency obligations in Article 50 of the EU AI Act became applicable, and two things started at once: the side that builds the models has to mark what it generates so machines can detect it, and the side that publishes content — you — has to disclose, in certain cases, that a passage was produced by AI.
The technical half was settled a month before the deadline. On 24 July Google signed the EU’s Code of Practice on Transparency of AI-Generated Content and opened up its SynthID watermarking to Apple, OpenAI, NVIDIA, ElevenLabs and Kakao so the marks would be interoperable across labs. The Code comes down to two mechanisms: signed metadata attached to a file, like a birth certificate for an image, and an imperceptible watermark written into the content itself.
The text half was published this week. Anthropic’s documentation treats 2 August as the trigger and is specific about what follows: new models carry, from launch day, a watermark woven into the generated text that no reader can see, one that survives copy-paste and may persist through light editing. Files (.png, .jpg, .svg) get a C2PA signature instead. The same page is candid about two limits — no mark detected doesn’t mean it wasn’t AI-written (short passages, rewrites and format conversions can all wash the signal out), and a mark detected doesn’t mean that lab wrote it (translate or summarise someone else’s work and the mark comes along). How anyone outside the company actually checks, the documentation says, will follow in a technical write-up.
Why it matters to you
Clear one thing out of the way first: these marks exist for regulatory detectability, and nothing in the public documentation ties them to search ranking. No engine pushes your pages down because your text carries a watermark.
Don’t read that as “nothing to worry about”, though. “Your ranking holds” and “the AI still knows you are there” are two different things, and the second one is what this piece is actually about.
The clause that reaches you is a different one. Article 50(4) says that when you publish AI-generated text to inform the public on matters of public interest, you must disclose that it was artificially generated or manipulated. Before you go and staple “written by AI” to the top of every post, read the carve-out attached to it, because the carve-out is the whole story:
The obligation does not apply where the content has undergone human review or editorial control and a natural or legal person holds editorial responsibility for its publication.
The law draws a line straight through the middle of content operations. On one side, AI drafts go out as they came; you owe the reader a label. On the other, a person has genuinely read it, changed it, and holds responsibility for publishing it; you owe nothing. The text does not require that person’s name to be visible, but you do need a trace showing someone was accountable, and in practice a byline is the cheapest way to leave one. Both sides drafted with AI. They do not look remotely alike to a reader.
One more thing that should let you sleep tonight: the European Commission only adopted its guidance on 20 July and the obligation itself only started on 2 August, so content generated and published before 2 August does not have to be labelled retroactively. Nobody is asking you to go back through three years of posts.
Get the penalties right, in both directions. Breaching Article 50 carries fines up to €15 million or 3% of worldwide annual turnover, whichever is higher — but it is EU law, and it reaches you only if you place services on the EU market or your output is used inside the EU. Most brands outside Europe will not be chased by that number this year; note, though, that the test written into the law is where the output lands, not where the company is registered. If you publish in English, do the maths yourself.
Two other pressures arrive sooner. Platform rules are one: English Wikipedia barred LLMs from generating or rewriting article text back in March (copyediting your own prose and assisted translation still stand), on the grounds that verifiability cannot keep up with generation: writing takes seconds, checking every claim back to a source takes hours. The closer one is detection itself — once the tools ship, as the law requires them to, your clients, your competitors and whoever inherits your website next can run them across everything you have ever published. On that day you will want something to show.
There’s a longer-range possibility that should worry anyone doing GEO more than the law does. What happens when training data gets polluted with AI’s own output was tested in a 2024 Nature paper: feed model-generated content back into training repeatedly and the model degrades, with the tails of the original distribution disappearing entirely. The authors’ conclusion was that synthetic data isn’t unusable — but the filtering has to be taken seriously. Until now nobody had a reliable way to pick AI text out of a corpus. Watermarks make that technically possible.
Nobody has said they intend to use it that way, and the step may never be taken. The cost is still worth holding in mind: not a few ranking positions, but your content getting scrubbed out of the next generation of models — the engine can’t recall you, because it never learned you.
Do you need to act now
Two things you can skip. Don’t switch off AI-assisted writing; no public document supports the idea that using it earns a search penalty, and if the corpus layer is what worries you, switching it off is not the fix either. What gets filtered out is content nobody stands behind, not content that was drafted with AI. And don’t go buying one of those “98% accurate AI detectors” to audit yourself — that class of tool guesses at writing style, which has nothing to do with the watermarks described here, and running it will hand you a meaningless number.
What does deserve a decision is deeply unglamorous: name a person in your content process who genuinely reads the draft, edits it, and is willing to have their name on it — and leave a trace that the review happened. That single step covers three things at once: the condition for the legal carve-out, the bar platforms like Wikipedia now set, and whatever trust your readers still extend to your site. It runs on the same logic as the provenance layer the industry started building last year: content that can account for its own origin.
As for when detection ships, or whether any engine will do something with the signal, nobody can give you a date — anyone promising you a timeline is selling something, and that applies to us too.
The direction, though, is no longer a guess. The value of content is moving from who can produce it to who stands behind it. The first is available to everyone as of this year. The second cannot be faked.
If you’re not sure which side of that line your own content process sits on, we can take a look first.