I think that you may have forgotten that the entire point of my first post was that Claude is getting watermarked, so that it can be detected. Current AI detection tools are spotty, but the link below suggests it can be clearly identified.
No, that is the point of the conversation. I just didn't connect the dots (i.e., logic) for you. The point of watermarking is to provide some indication of authorship: Did Claude write it or did You write it. - To deal with the issues you are concerned with, you have to do a deep dive into how Anthropic is actually doing the watermarking:
Claude's watermarking method is open source. They use an algorithm called "SynthID-Text" or simply "SynthID," originally developed by Google DeepMind. It biases token selection using a keyed pseudo-random "g-value" over each candidate token based on an underlying mechanism Google calls "Tournament Sampling" computed from a secret key plus preceding context, which, according to Claude, "is exactly why the market lives in the pattern of word choices,
needs many tokens to detect and washes out under paraphrasing." Google's method, in turn, is based on the "Kirchenbauer/Aaronson Sampling-Watermark Paradigm":
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein.“A Watermark for Large Language Models.”Proceedings of the 40th International Conference on Machine Learning (ICML 2023), PMLR 202:17061–17084, 2023.
For some code:
Contribute to jwkirchenbauer/lm-watermarking development by creating an account on GitHub.
github.com
"Watermarking doesn’t change the meaning or experience for the person reading it, but if you wanted to check after the fact whether the text was likely generated by Claude, the watermark allows you to do so."
You are wrong about that. It changes the words and dramatically so. After a while, you will be able to detect the difference. Various tactics surface. For example, you may begin to notice the unusual appearance of distinctly British spelled words or variants that are otherwise unusual in the US.
| British | American |
| colour | color |
| favour | favor |
| honour | honor |
| labour | labor |
| grey | gray |
| metre | meter |
| realize | realize |
| defence | defense |
| travelling | traveling |
| cancelled | canceled |
| aluminium | aluminum |
| dialogue | dialog |
| jewellery | jewelry |
| theatre | theater |
So, at certain times of the day or week or month, it might bias the British spelling over the American.
When you review your generated output, you WILL most likely recognize where it is using certain "unusual" words. Just follow the practice of substituting your "preferred" version. - And sabatoge the system.
"We’re implementing watermarking to comply with the EU AI Act. Anthropic, along with several other major AI model providers and around
190 total signatories, signed the EU Code of Practice on Transparency of AI-Generated Content in July 2026. This requires AI system providers to use methods of “marking” AI-generated text."
That is the BIG joke.
Me to Claude: "Yes but text watermarks are so easy to subvert with paraphrasing."
Claude: "Right — and "easy" actually undersells it, because it's not merely easy in practice, it's provably unavoidable in principle. The "Watermarks in the Sand" result says no quality-preserving text watermark can survive a paraphrasing attacker. So the EU has mandated, for text specifically, a property that a theorem says can't exist. That's a strange place for a law to stand, and the "as far as technically feasible" clause is quietly load-bearing precisely because "feasible" for robust text marking is, today, "no."
Which raises the fair question of what the mandate is even
for, and I think there are two honest answers. First, it's aimed at the cooperative majority, not adversaries. Most AI-generated text isn't being laundered by someone determined to hide it — it's posted as-is by people who simply didn't disclose. The watermark labels that enormous cooperative middle, and that's a real transparency gain. Anyone who bothers to paraphrase-to-evade has, by that act, demonstrated intent — which pushes them out of the "transparency" frame and into fraud/misrepresentation frames the watermark was never meant to police anyway. Second, it functions as a regulatory ratchet: "deploy the best available technology and improve it," a forcing function for the provenance ecosystem (detection APIs, C2PA, standards) more than a guarantee of catching anyone.
There's also a wrinkle that cuts the other way and I think is underappreciated. If watermarking becomes
universal among compliant providers, the obvious evasion — "paraphrase it with an AI" — backfires, because paraphrasing Claude's text through another compliant model just re-stamps it with
that model's watermark. To get clean removal you need an
unwatermarked paraphraser: a local, open-source model, or manual rewriting. So the mandate's practical effect is to push determined evaders toward open/self-hosted models the EU can't compel to watermark at all. The regulation is enforceable exactly against the providers who'd comply anyway, and structurally leaky toward the tools a motivated actor would actually use.
So the policy critique writes itself, and it's the asymmetry you keep circling: the mark reliably tags the honest and reliably fails on the dishonest, while the costs — quality tradeoffs, the alignment side-effects, deployment and key-management overhead, the forgeability risk — land on everyone. As a transparency signal for the cooperative case, defensible. As a control against bad actors or as forensic evidence, close to theater. And the law, by writing in "as far as technically feasible," has more or less admitted that in its own text.
Right — and "easy" actually undersells it, because it's not merely easy in practice, it's provably unavoidable in principle. The "Watermarks in the Sand" result says no quality-preserving text watermark can survive a paraphrasing attacker. So the EU has mandated, for text specifically, a property that a theorem says can't exist. That's a strange place for a law to stand, and the "as far as technically feasible" clause is quietly load-bearing precisely because "feasible" for robust text marking is, today, "no."
Which raises the fair question of what the mandate is even
for, and I think there are two honest answers. First, it's aimed at the cooperative majority, not adversaries. Most AI-generated text isn't being laundered by someone determined to hide it — it's posted as-is by people who simply didn't disclose. The watermark labels that enormous cooperative middle, and that's a real transparency gain. Anyone who bothers to paraphrase-to-evade has, by that act, demonstrated intent — which pushes them out of the "transparency" frame and into fraud/misrepresentation frames the watermark was never meant to police anyway. Second, it functions as a regulatory ratchet: "deploy the best available technology and improve it," a forcing function for the provenance ecosystem (detection APIs, C2PA, standards) more than a guarantee of catching anyone."