How to use this check
Scan the text for observable characters
The scanner names every invisible or control character with its code point, count, and positions.
Read the finding as what it is
A hidden character tells you something about the text's copy-and-paste history, not about which system produced it.
Do not convert it into an authorship claim
There is no threshold of invisible characters that indicates AI authorship, because the mechanisms are unconnected.
Hidden characters are not a Claude detector
Zero-width characters, directional controls, non-breaking spaces, and soft hyphens appear in text from a long list of entirely ordinary workflows. Copying from PDFs, document editors, messaging apps, content management systems, and websites all introduce them.
The watermarking concept discussed for language models concerns statistical patterns in word selection, not the insertion of secret Unicode characters. These are different mechanisms operating at different levels of the text, and finding one tells you nothing about the presence of the other.
So a text full of zero-width spaces is most likely a text that was pasted out of a PDF. A text with none of them is most likely a text that was typed or pasted from a plain editor. Neither observation identifies the author.
How statistical text watermarking is supposed to work
The published research approach works at the token level. When a model generates text it chooses each next token from a probability distribution. A watermarking scheme uses a secret key to partition the vocabulary and nudges the sampler toward one partition, in a way that barely affects quality but leaves a statistical bias.
A detector holding the same key can then measure how often the generated text falls into the favoured partition and compute how unlikely that pattern would be by chance. Given enough text, this produces a genuine statistical test.
Two properties of this design matter for anyone hoping to check text. First, detection requires the key — it is not something an outside party can reconstruct. Second, it requires a meaningful volume of text, because the signal is statistical and short passages do not carry enough evidence.
Why the signal degrades under normal use
Paraphrasing substitutes tokens and destroys the pattern. Translation destroys it completely. Heavy editing dilutes it. Mixing generated and human-written passages leaves a weak signal spread across text that also contains unwatermarked material.
Short text is a separate problem. A statistical test over a few sentences has very little power, which means a real watermark can easily fail to reach significance in exactly the length of text people most often want to check.
These are not implementation flaws; they are inherent to a statistical approach. Any detector has to trade false positives against false negatives, and the consequences of a false positive — accusing someone of undisclosed AI use — are severe enough that a cautious threshold is the right choice, which in turn means missing real cases.
What Anthropic has actually said
Anthropic's published transparency material describes continuing work on provenance and watermarking rather than an available detection product. There is no official, supported Claude text detector that a third party can call.
Until such a mechanism exists and is documented, any site claiming to detect Claude-authored text is doing something else — usually a generic classifier trained to guess from style — and presenting the output as if it were a watermark measurement.
That distinction matters a great deal when the result is used to make a decision about a person. A generic style classifier has a false-positive rate that falls hardest on writers whose prose is unusually clean, formal, or non-native.
What to do if you need to make a decision
If you are assessing student or employee work, the technically honest position is that no available tool gives you a reliable verdict on a specific piece of text. Process evidence is stronger and fairer: drafts, version history, an ability to discuss the reasoning, or a short supervised writing sample.
If you are the one being accused, ask which specific mechanism the tool claims to measure and what its published false-positive rate is. A tool that cannot answer either question is not evidence.
If you are publishing, disclosure is far more robust than detection. The direction of regulation — including the EU's transparency obligations — is toward declaring AI involvement rather than trying to detect it after the fact.
Frequently asked questions
Can this tool detect if text was written by Claude?
No, and no public tool currently can. There is no official Anthropic detector available to third parties, and the statistical approach described in the research requires a secret key that only the provider holds.
Does Claude insert hidden characters into its output?
No. The invisible characters this scanner finds come from copying between applications, PDF extraction, and encoding conversion. Language models generate ordinary text.
Why do some sites claim to detect Claude text?
They are almost always running a generic style classifier and describing it as watermark detection. Ask what mechanism they measure and what their published false-positive rate is.
Would a Claude watermark survive editing?
A statistical watermark weakens under paraphrasing, shortening, mixing with human text, and disappears under translation. Even with the correct detector, edited text is hard to assess.
What is a fair way to handle a suspected AI submission?
Process evidence rather than detector output: drafts and version history, a conversation about the reasoning, or a supervised writing sample. Detector scores are not reliable enough to support a decision about a person.
Primary sources
The technical claims on this page follow the published specifications below rather than our own assertions.
- Anthropic transparency commitments
Anthropic's published position on provenance and watermarking work.
- European Commission AI transparency guidance
The EU framework driving machine-readable marking obligations for synthetic content.
- Unicode Character Database (UAX #44)
Defines the general categories, including the Cf format characters this scanner names.