The AI Watermark Flags the Wrong People. Here’s Why.

By Lynn Räbsamen, CFA | Advisory Board Member, CFA Institute | Author, Artificial Stupelligence

Anthropic confirmed that Claude models now embed an imperceptible watermark directly into generated text, across the API, the chatbot, Claude Code, and the major clouds, worldwide rather than only in Europe. It is a compliance move under Article 50 of the EU AI Act which went live since August 2, 2026.

What the AI watermark actually is

The mechanism is worth understanding, because almost every commentary gets it wrong. The implementation adapts Google DeepMind’s SynthID-Text, published in Nature in 2024, and it changes only the source of randomness used to select the next word.

A model writes one word at a time. At most steps, several candidates fit equally well. Cold and overcast, or cold and grey. Nobody notices either way, and normally a random number decides. Watermarking swaps that random number for a function driven by a secret key. Any single word looks unremarkable. Several hundred words later, the pattern is measurable.

The watermark is not attached to the text. The watermark is the text.

Which is why it survives copy and paste. You are not carrying a hidden tag. You are carrying the words, and the words are the tag.

Why this is not GPTZero

This distinction matters more than the news itself, and it is being flattened in every LinkedIn thread I have read.

ZeroGPT, GPTZero and their competitors are guessers. They have no access to Anthropic’s keys. They scan for patterns typical of AI text: telltale phrasing, overused words, uniform rhythm. A style classifier, trained on vibes and hardened by nothing. This is why they keep flagging non-native English speakers, and why people defeat them by adding typos.

Watermark detection asks a narrower question. Does the statistical skew in this passage match what our key would produce. That is a fundamentally different approach and should prove more reliable in practice.

Reliable at answering a different question, though. And a smaller one.

It also matters where you would run the check. As of writing, Anthropic has not shipped a detection API, a web tool, or a technical spec. Detection tooling is announced and in development. The mark is live now. That gap is the interesting part: a partially understood signal is already circulating in documents before anyone has agreed what it proves.

What the mark cannot tell you

Anthropic is admirably direct about the limits, in language nobody downstream will read.

A detected mark indicates content may have been processed by Claude. Claude may not be the original author, because people routinely use it to proofread, translate, summarize, and convert files. The output can carry a mark even when the underlying ideas and text came from a person.

Then the reverse. Absence of a mark does not mean content was not AI-generated: the text may have been heavily edited, paraphrased, translated, mixed into other writing, or simply too short to carry a reliable signal.

So the analyst who wrote her own memo and ran it through Claude for grammar produces a marked document. The analyst who generated the whole thing and then asked a different model to rewrite it produces a clean one.

One of them did the work. The mark flags the other one.

The Big 4 problem, restated

Somewhere in a 237-page report delivered to the Australian government sits a quote from a Federal Court judge. The judge never said it.

The report cost 440,000 Australian dollars. Deloitte Australia agreed to partially refund it. The revised version, published after the fact, disclosed that Azure OpenAI had been used in the writing, and removed the fabricated judicial quote along with references to research papers that do not exist. A University of Sydney researcher found the errors and told the media. Not a detector. A person, reading footnotes.

Now put the AI watermark back into that 237-page report and ask what it would have changed.

It would have confirmed that a model touched the draft. Deloitte disclosed exactly that in the revised version, after the fact, and the disclosure did not fix anything.

The problem was never that a language model was in the room.

The problem was that a fabricated quote from a Federal Court judgment traveled from a chat window into a government deliverable without anyone checking whether the judge had said it.

I wrote about the Big 4 and the Act’s disclosure regime here. The pattern has not changed. Firms are getting better at declaring AI involvement and no better at verifying AI output. Those are separate disciplines, and only one of them is expensive.

Disclosure is cheap. Verification takes a qualified person, billable hours, and a willingness to find something wrong late in the process.

The question worth asking

Drafting with AI is not the problem. Running a paragraph through a model to tighten it is not the problem. Translating a client letter into German is not the problem. If we are honest, most of the profession has been doing some version of this for two and a half years, and the output is frequently better than what preceded it.

The problem is a document nobody stands behind.

The only question that has ever mattered is whether a competent human read this, checked the load-bearing claims, and will own the consequences if it is wrong.

That question predates the AI watermark by about 400 years.

It is the same question we ask of a junior analyst’s spreadsheet, an outsourced due diligence pack, and a research note with someone else’s name on the cover.

Provenance was never the control. Review was. The watermark will not answer it.

Neither will ZeroGPT, an honor code hearing, or an em dash count. And there is a real risk in the interim, which is that a probabilistic signal about processing gets treated as proof of authorship by people who need a fast answer and will not read the limitations page.

That is how you end up sanctioning the analyst who used spellcheck while the person who published a hallucinated citation walks free with clean text.

Sign your work. Check it first. The watermark is a fingerprint on a door handle, and we are about to spend a great deal of energy dusting for it while the actual question stands in the corner, unasked.

This article was partially drafted by AI and reviewed by a human. There may be a watermark. There was definitely a human.


For more insights about what AI can or cannot do, check out my book “Artificial Stupelligence: The Hilarious Truth About AI”.

Subscribe here to be the first to receive my insights.


Discover more from Lynn Raebsamen, CFA

Subscribe to get the latest posts sent to your email.

Love this content? Get updates in your inbox.

Subscribe now to keep reading and get access to the full archive.

Continue reading