The direct answer
Anthropic's public transparency material describes ongoing work on content provenance and watermarking. It does not publish an available detection endpoint or tool that a third party can use to test whether a given passage came from Claude.
That means two things at once, and both need stating. Nobody outside Anthropic can currently verify a Claude text watermark. And the absence of a public detector is not proof that no marking exists — it is simply the state of what has been made available.
Anyone telling you they can detect Claude-authored text today is describing a different thing: a classifier that guesses from writing style. That is not watermark verification, and it should not be presented as one.
Statistical watermarking is about token choices
Text watermarking research generally works by nudging the probability of selected next tokens according to a secret key, creating a pattern that a compatible detector can test for across a sufficient volume of text. The output reads normally; the bias is only detectable statistically.
This is fundamentally different from inserting zero-width spaces or unusual punctuation. Those are visible to anyone who looks and removable by anyone who cares. A statistical watermark cannot be seen by inspection at all, and cannot be removed by find-and-replace.
It also cannot be checked without the key. The whole security property of the scheme depends on the partition being unguessable, which is precisely why third-party detection is not possible by design rather than by oversight.
Why hidden Unicode became the popular myth
The claim spreads because it is checkable. Anyone can paste text into a tool, see a zero-width space, and feel they have found something. The mechanism is satisfying in a way that a statistical argument is not.
But the causal story does not hold. Those characters arrive from PDF extraction, word processors, web page copy, messaging apps, and encoding conversion. Text that has been through a document pipeline is full of them regardless of who wrote it, and text typed directly into a plain editor has none regardless of who wrote it.
The test therefore measures the copy-and-paste history of a passage, which is real information but has nothing to do with authorship.
What would need to exist for a reliable check
Three things, none of which are currently in place for third parties.
- A published, documented marking scheme, so the mechanism being tested is known rather than assumed.
- An official detection interface operated by the provider who holds the key.
- Published performance characteristics — a false-positive rate, a minimum text length, and behaviour under paraphrasing and translation — so results can be interpreted rather than trusted blindly.
What to rely on instead
For images and video, signed provenance is genuinely verifiable today. C2PA Content Credentials can be validated by anyone with public tooling, which is why this site checks them and reports the result honestly.
For text, disclosure beats detection. The regulatory direction — including the EU's transparency obligations for synthetic content — is toward declaring AI involvement rather than trying to reconstruct it afterwards, precisely because reliable after-the-fact detection of text does not exist.
For assessment specifically, process evidence is both more reliable and fairer than any detector: drafts, revision history, and the ability to discuss the work.
Frequently asked questions
Does Claude add a watermark to its text output?
Anthropic describes continuing work on watermarking and provenance but has not published an available detection mechanism. No third party can currently verify whether a given passage carries a mark.
Do hidden Unicode characters prove Claude wrote something?
No. Those characters come from PDFs, editors, web pages, and encoding conversion. They reflect how a passage was copied around, not which system produced it.
Are third-party Claude detectors reliable?
They are running style classifiers, not watermark verification. Their false positives fall hardest on formal, clean, or non-native writing, which makes them unsuitable for decisions about individuals.
Would a watermark survive if I edited the text?
Statistical watermarks weaken under paraphrasing and shortening, and are destroyed by translation. Even the provider's own detector would struggle with heavily edited text.
Is there any way to prove text was written by a human?
Not through analysis of the finished text. Process evidence — drafts, version history, and an ability to discuss the reasoning — is the only practical demonstration.
Primary sources
The technical claims on this page follow the published specifications below rather than our own assertions.
- Anthropic transparency commitments
Anthropic's published position on provenance and watermarking work.
- European Commission AI transparency guidance
The EU framework driving machine-readable marking obligations for synthetic content.
- C2PA technical specification 2.3
The normative definition of manifests, claims, assertions, hard bindings, and validation states.