• Welcome to AppraisersForum.com, the premier online  community for the discussion of real estate appraisal. Register a free account to be able to post and unlock additional forums and features.

A.I. Stuff

No, that is the point of the conversation. I just didn't connect the dots (i.e., logic) for you. The point of watermarking is to provide some indication of authorship: Did Claude write it or did You write it. - To deal with the issues you are concerned with, you have to do a deep dive into how Anthropic is actually doing the watermarking:
Claude's watermarking method is open source. They use an algorithm called "SynthID-Text" or simply "SynthID," originally developed by Google DeepMind. It biases token selection using a keyed pseudo-random "g-value" over each candidate token based on an underlying mechanism Google calls "Tournament Sampling" computed from a secret key plus preceding context, which, according to Claude, "is exactly why the market lives in the pattern of word choices, needs many tokens to detect and washes out under paraphrasing." Google's method, in turn, is based on the "Kirchenbauer/Aaronson Sampling-Watermark Paradigm":

John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein.“A Watermark for Large Language Models.”Proceedings of the 40th International Conference on Machine Learning (ICML 2023), PMLR 202:17061–17084, 2023.

For some code:


You are wrong about that. It changes the words and dramatically so. After a while, you will be able to detect the difference. Various tactics surface. For example, you may begin to notice the unusual appearance of distinctly British spelled words or variants that are otherwise unusual in the US.

BritishAmerican
colourcolor
favourfavor
honourhonor
labourlabor
greygray
metremeter
realizerealize
defencedefense
travellingtraveling
cancelledcanceled
aluminiumaluminum
dialoguedialog
jewelleryjewelry
theatretheater

So, at certain times of the day or week or month, it might bias the British spelling over the American.

When you review your generated output, you WILL most likely recognize where it is using certain "unusual" words. Just follow the practice of substituting your "preferred" version. - And sabatoge the system.


That is the BIG joke.

Me to Claude: "Yes but text watermarks are so easy to subvert with paraphrasing."

Claude: "Right — and "easy" actually undersells it, because it's not merely easy in practice, it's provably unavoidable in principle. The "Watermarks in the Sand" result says no quality-preserving text watermark can survive a paraphrasing attacker. So the EU has mandated, for text specifically, a property that a theorem says can't exist. That's a strange place for a law to stand, and the "as far as technically feasible" clause is quietly load-bearing precisely because "feasible" for robust text marking is, today, "no."

Which raises the fair question of what the mandate is even for, and I think there are two honest answers. First, it's aimed at the cooperative majority, not adversaries. Most AI-generated text isn't being laundered by someone determined to hide it — it's posted as-is by people who simply didn't disclose. The watermark labels that enormous cooperative middle, and that's a real transparency gain. Anyone who bothers to paraphrase-to-evade has, by that act, demonstrated intent — which pushes them out of the "transparency" frame and into fraud/misrepresentation frames the watermark was never meant to police anyway. Second, it functions as a regulatory ratchet: "deploy the best available technology and improve it," a forcing function for the provenance ecosystem (detection APIs, C2PA, standards) more than a guarantee of catching anyone.

There's also a wrinkle that cuts the other way and I think is underappreciated. If watermarking becomes universal among compliant providers, the obvious evasion — "paraphrase it with an AI" — backfires, because paraphrasing Claude's text through another compliant model just re-stamps it with that model's watermark. To get clean removal you need an unwatermarked paraphraser: a local, open-source model, or manual rewriting. So the mandate's practical effect is to push determined evaders toward open/self-hosted models the EU can't compel to watermark at all. The regulation is enforceable exactly against the providers who'd comply anyway, and structurally leaky toward the tools a motivated actor would actually use.

So the policy critique writes itself, and it's the asymmetry you keep circling: the mark reliably tags the honest and reliably fails on the dishonest, while the costs — quality tradeoffs, the alignment side-effects, deployment and key-management overhead, the forgeability risk — land on everyone. As a transparency signal for the cooperative case, defensible. As a control against bad actors or as forensic evidence, close to theater. And the law, by writing in "as far as technically feasible," has more or less admitted that in its own text.
Right — and "easy" actually undersells it, because it's not merely easy in practice, it's provably unavoidable in principle. The "Watermarks in the Sand" result says no quality-preserving text watermark can survive a paraphrasing attacker. So the EU has mandated, for text specifically, a property that a theorem says can't exist. That's a strange place for a law to stand, and the "as far as technically feasible" clause is quietly load-bearing precisely because "feasible" for robust text marking is, today, "no."

Which raises the fair question of what the mandate is even for, and I think there are two honest answers. First, it's aimed at the cooperative majority, not adversaries. Most AI-generated text isn't being laundered by someone determined to hide it — it's posted as-is by people who simply didn't disclose. The watermark labels that enormous cooperative middle, and that's a real transparency gain. Anyone who bothers to paraphrase-to-evade has, by that act, demonstrated intent — which pushes them out of the "transparency" frame and into fraud/misrepresentation frames the watermark was never meant to police anyway. Second, it functions as a regulatory ratchet: "deploy the best available technology and improve it," a forcing function for the provenance ecosystem (detection APIs, C2PA, standards) more than a guarantee of catching anyone."

If you are a little mixed up after carefully reading this, please note that Anthropic is trying to save itself from drowning, it's being hammered on different sides in a maybe possibly sinking ship. Will it survive?

Grok says, quite the contrary, IT IS NOT WATERMARKING TEXT!!! Strange, the EU mandates it.

Here is Grok's official, if dubious, stance:

xAI / Grok’s position​

"xAI did not sign the voluntary Code of Practice. Official statements indicate that Grok text outputs carry no invisible statistical watermarks, fingerprints, or tracking signals, and that they meet applicable EU rules through alternative means that keep the text clean. Visible marks are applied to images and videos generated via Imagine features.

Bottom line: The EU requires machine-readable marking and detectability of synthetic text from providers serving the EU market, but it does not specifically require the Kirchenbauer/Aaronson-style sampling watermark. Other technical approaches are allowed if they satisfy the statutory criteria of effectiveness, robustness, etc."
 
The whole AI sphere moving to the EU standard is the gross part lol

Yeah I don't get it either. I mean the assumption is AI is assisting everyone, everywhere, so why the need to watermark?

Maybe it's more to do with control via EU Copyright law?
 
AI may seem free or cheap now but AI companies will be charging users more in the future as we use it more and find it useful.
I heard Google phones have new AI software and Apple hasn't kept up. If iphone doesn't have the AI capabilites like Google, it will be the demise of the iphone.
 
There are a number of problems identifying, with any degree of certainty, machine generated text:

1. To get a 99% probability of identifying medium-entropy (complexity) English text with at least 99% probability, you need about 400 words. That is assuming no modification of the output text, no paraphrasing.

2. However, if we are talking about identifying machine-generated text within a professional niche like appraisal, which has a rather constrained vocabulary:

"Real estate reports (appraisals, BPOs, CMAs, market analyses, etc.) are usually long enough that length itself is not the limiting factor. A typical narrative section easily exceeds 300–400 words. The difficulties come from other characteristics of the work:
  • Lower entropy in key sections. Property descriptions, comparable sales write-ups, standard disclaimer language, regulatory citations, and numerical tables leave the model with fewer free choices. The statistical watermark is weaker precisely where the text is most constrained and formulaic.
  • Heavy human editing is normal. Professional reports almost always go through substantial revision, fact-checking, and rephrasing. Even moderate rewriting dilutes or erases the detectable signal. After a few rounds of ordinary professional editing, confident detection often becomes unreliable.
  • Mixed provenance. Many workflows start with an AI draft, then add human analysis, local market knowledge, photos, and calculations. The final document is a hybrid; a watermark detector may flag only portions or return an ambiguous result.
  • Volume and consistency pressure. Generating reports day after day with the same commercial model means every output carries the mark. If a client, reviewer, lender, or regulator later runs a detector and finds the mark, it can raise questions even when the final work product meets professional standards. Conversely, if you route the text through another model or heavy post-processing to remove the mark, you add friction and cost to an already high-volume process.
  • Open or non-compliant models. Using open-weight models (or any system that does not watermark) avoids the issue entirely, but then you lose whatever quality or convenience advantage the major commercial systems provide.
In short, the watermarking regime works better against casual, lightly edited, high-volume synthetic content than against carefully curated professional documents that mix AI assistance with human expertise and domain-specific constraints. For day-to-day real estate report generation, the mark is often either easy to dilute through normal editing or present enough to create downstream friction depending on who is checking. " (Grok)

The reality is: Some kinds of text can be reliably watermarked, especially creative text that is categorically different. Other text, think professional real estate, is mundane, boring, repetitious, and you're going to have a hard time telling the difference between machine-generated and human-written.

On the other hand, for those fancy writers out there, anyone who gets their hands on your creative output and submits it to Claude for analysis might be surprised at what Claud has to say. It can do a very thorough job of analyzing the method behind the madness, - so to say. Well, I won't go into it any further.

I say, whatever some AI generates for you, always review it and be very ready to change anything at all that doesn't suit you. It is good practice anyway to thoroughly review AI generated documents. I could tell you some stories. But I will save them.
 
When Fernando types his comments in reports (let alone in AF), y'all know it wasn't AI generated.
 
A good way to humanize AI text is to remove all em dashes, (ii) citations no one uses and get to the point with facts then conclusion. I used gpt today to synthesize my ramblings which it did far more than I ever would while keeping the intent throughout. The various checkers I started out with started at 94% AI, then 60ish finally at 5% when I shortened a 4 page doc to 1 instructing it to get to the effing point


aibot.png
 
when I have a beer or 9 I like to query goog about appraisal services from various locations through a free vpn. tonight was an epic night

1787200083059.png
 
Just occurred to me that 20 years ago I was kicked off this forum for discussing ways to subvert some security algorithm. Then I woke up this morning dreaming I was being investigated by the Appraisal Institute for discussing ways to subvert this watermarking method. Seriously, I wonder ....
 
Find a Real Estate Appraiser - Enter Zip Code

Copyright © 2000-, AppraisersForum.com, All Rights Reserved
AppraisersForum.com is proudly hosted by the folks at
AppraiserSites.com
Back
Top