- Joined
- Jun 27, 2017
- Professional Status
- Certified General Appraiser
- State
- California
No, that is the point of the conversation. I just didn't connect the dots (i.e., logic) for you. The point of watermarking is to provide some indication of authorship: Did Claude write it or did You write it. - To deal with the issues you are concerned with, you have to do a deep dive into how Anthropic is actually doing the watermarking:
Claude's watermarking method is open source. They use an algorithm called "SynthID-Text" or simply "SynthID," originally developed by Google DeepMind. It biases token selection using a keyed pseudo-random "g-value" over each candidate token based on an underlying mechanism Google calls "Tournament Sampling" computed from a secret key plus preceding context, which, according to Claude, "is exactly why the market lives in the pattern of word choices, needs many tokens to detect and washes out under paraphrasing." Google's method, in turn, is based on the "Kirchenbauer/Aaronson Sampling-Watermark Paradigm":
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein.“A Watermark for Large Language Models.”Proceedings of the 40th International Conference on Machine Learning (ICML 2023), PMLR 202:17061–17084, 2023.
For some code:
![]()
GitHub - jwkirchenbauer/lm-watermarking
Contribute to jwkirchenbauer/lm-watermarking development by creating an account on GitHub.github.com
You are wrong about that. It changes the words and dramatically so. After a while, you will be able to detect the difference. Various tactics surface. For example, you may begin to notice the unusual appearance of distinctly British spelled words or variants that are otherwise unusual in the US.
British American colour color favour favor honour honor labour labor grey gray metre meter realize realize defence defense travelling traveling cancelled canceled aluminium aluminum dialogue dialog jewellery jewelry theatre theater
So, at certain times of the day or week or month, it might bias the British spelling over the American.
When you review your generated output, you WILL most likely recognize where it is using certain "unusual" words. Just follow the practice of substituting your "preferred" version. - And sabatoge the system.
That is the BIG joke.
Me to Claude: "Yes but text watermarks are so easy to subvert with paraphrasing."
Claude: "Right — and "easy" actually undersells it, because it's not merely easy in practice, it's provably unavoidable in principle. The "Watermarks in the Sand" result says no quality-preserving text watermark can survive a paraphrasing attacker. So the EU has mandated, for text specifically, a property that a theorem says can't exist. That's a strange place for a law to stand, and the "as far as technically feasible" clause is quietly load-bearing precisely because "feasible" for robust text marking is, today, "no."
Which raises the fair question of what the mandate is even for, and I think there are two honest answers. First, it's aimed at the cooperative majority, not adversaries. Most AI-generated text isn't being laundered by someone determined to hide it — it's posted as-is by people who simply didn't disclose. The watermark labels that enormous cooperative middle, and that's a real transparency gain. Anyone who bothers to paraphrase-to-evade has, by that act, demonstrated intent — which pushes them out of the "transparency" frame and into fraud/misrepresentation frames the watermark was never meant to police anyway. Second, it functions as a regulatory ratchet: "deploy the best available technology and improve it," a forcing function for the provenance ecosystem (detection APIs, C2PA, standards) more than a guarantee of catching anyone.
There's also a wrinkle that cuts the other way and I think is underappreciated. If watermarking becomes universal among compliant providers, the obvious evasion — "paraphrase it with an AI" — backfires, because paraphrasing Claude's text through another compliant model just re-stamps it with that model's watermark. To get clean removal you need an unwatermarked paraphraser: a local, open-source model, or manual rewriting. So the mandate's practical effect is to push determined evaders toward open/self-hosted models the EU can't compel to watermark at all. The regulation is enforceable exactly against the providers who'd comply anyway, and structurally leaky toward the tools a motivated actor would actually use.
So the policy critique writes itself, and it's the asymmetry you keep circling: the mark reliably tags the honest and reliably fails on the dishonest, while the costs — quality tradeoffs, the alignment side-effects, deployment and key-management overhead, the forgeability risk — land on everyone. As a transparency signal for the cooperative case, defensible. As a control against bad actors or as forensic evidence, close to theater. And the law, by writing in "as far as technically feasible," has more or less admitted that in its own text.
Right — and "easy" actually undersells it, because it's not merely easy in practice, it's provably unavoidable in principle. The "Watermarks in the Sand" result says no quality-preserving text watermark can survive a paraphrasing attacker. So the EU has mandated, for text specifically, a property that a theorem says can't exist. That's a strange place for a law to stand, and the "as far as technically feasible" clause is quietly load-bearing precisely because "feasible" for robust text marking is, today, "no."
Which raises the fair question of what the mandate is even for, and I think there are two honest answers. First, it's aimed at the cooperative majority, not adversaries. Most AI-generated text isn't being laundered by someone determined to hide it — it's posted as-is by people who simply didn't disclose. The watermark labels that enormous cooperative middle, and that's a real transparency gain. Anyone who bothers to paraphrase-to-evade has, by that act, demonstrated intent — which pushes them out of the "transparency" frame and into fraud/misrepresentation frames the watermark was never meant to police anyway. Second, it functions as a regulatory ratchet: "deploy the best available technology and improve it," a forcing function for the provenance ecosystem (detection APIs, C2PA, standards) more than a guarantee of catching anyone."
If you are a little mixed up after carefully reading this, please note that Anthropic is trying to save itself from drowning, it's being hammered on different sides in a maybe possibly sinking ship. Will it survive?
Grok says, quite the contrary, IT IS NOT WATERMARKING TEXT!!! Strange, the EU mandates it.
Here is Grok's official, if dubious, stance:
xAI / Grok’s position
"xAI did not sign the voluntary Code of Practice. Official statements indicate that Grok text outputs carry no invisible statistical watermarks, fingerprints, or tracking signals, and that they meet applicable EU rules through alternative means that keep the text clean. Visible marks are applied to images and videos generated via Imagine features.Bottom line: The EU requires machine-readable marking and detectability of synthetic text from providers serving the EU market, but it does not specifically require the Kirchenbauer/Aaronson-style sampling watermark. Other technical approaches are allowed if they satisfy the statutory criteria of effectiveness, robustness, etc."


