How to use this check
Paste your text
Drop the text into the box. It stays in the page's memory and is never transmitted.
Run the scan
Each invisible or control character is grouped by code point with a count and the positions where it occurs.
Read the explanation for each finding
The report says what the character does and flags the ones that carry linguistic meaning.
Clean only what you chose
Safe mode removes just transport artifacts. Advanced mode lets you select individual categories.
What counts as a hidden character
The Unicode standard includes a category of format characters — general category Cf — that exist to control how surrounding text behaves rather than to be seen. They have no visual glyph of their own, which is exactly what makes them useful and exactly what makes them confusing when they appear unexpectedly.
This scanner also reports a handful of characters that are visible as whitespace but behave differently from an ordinary space, such as the non-breaking space and the narrow no-break space. These frequently cause search failures and layout surprises precisely because they look identical to something they are not.
- Zero-width space (U+200B) — a line-break opportunity with no width.
- Zero-width non-joiner (U+200C) and joiner (U+200D) — control whether adjacent letters connect; essential in Arabic, Persian, and Indic scripts, and used to build emoji sequences.
- Soft hyphen (U+00AD) — marks where a word may be hyphenated if it needs breaking.
- Byte order mark (U+FEFF) — a leftover encoding signal, usually a transport artifact.
- Non-breaking space (U+00A0) and narrow no-break space (U+202F) — prevent line breaks between words or before units.
- Bidirectional controls (U+202A–U+202E, U+2066–U+2069) — manage mixed left-to-right and right-to-left text.
- Word joiner (U+2060) — prevents a break without introducing a space.
Where these characters actually come from
The overwhelmingly common source is ordinary copying. PDF extraction inserts soft hyphens where the typesetter broke a word across lines. Word processors insert non-breaking spaces to keep a number attached to its unit. Web pages use zero-width spaces to allow long URLs to wrap. Messaging apps and CMS editors add byte order marks during encoding conversions.
In right-to-left and Indic text, joiners and directional marks are not artifacts at all — they are load-bearing. Removing a zero-width non-joiner from Persian text changes which letters connect, producing a word that is misspelled to a native reader while looking fine to someone who does not read the script.
Emoji are the other place where a joiner is doing essential work. A family emoji, a profession emoji, or a skin-tone variant is built by joining several code points with U+200D. Strip the joiners and one emoji becomes several unrelated ones.
The cases where hidden characters are a real problem
Search and comparison failures are the most common practical harm. A non-breaking space where a normal space belongs makes text unfindable by search, breaks exact-match filters in spreadsheets, and causes deduplication to miss obvious duplicates.
In source code and configuration files, an invisible character inside a string literal or an identifier produces errors that are effectively impossible to see in a diff. This is well documented as a code-review evasion technique, and bidirectional overrides in particular can make code display in an order that differs from how it executes.
In data pipelines, invisible characters in a key field silently break joins. The two values look identical on screen and compare as unequal, which produces a bug that survives a long time because nobody suspects the data itself.
What this scanner deliberately does not claim
Finding invisible characters in text does not indicate that the text was written by an AI system. This claim circulates widely and is wrong in both directions, alongside the other supposed writing tells.
Large language models produce ordinary Unicode text. They do not insert secret zero-width markers into their output as a signature. The characters this tool finds arrive through copying, editing, and format conversion — the same way they always have.
Conversely, a text with no hidden characters tells you nothing about its authorship either. This scanner reports observable characters, and that is the entire scope of what it can support.
How to decide what to remove
The safe default is to remove only characters that are almost always transport artifacts: the byte order mark and the soft hyphen. Neither carries meaning in normal running text, and both cause problems downstream — the cleaner removes exactly those two by default.
Everything else deserves a look at its context. If the text contains any Arabic, Persian, Urdu, Hebrew, Hindi, Bengali, or other script that uses joining behaviour, leave joiners and directional marks alone unless you can read the text and confirm they are wrong. If it contains emoji, leave U+200D alone.
If you are cleaning a key field, an identifier, or code, be much more aggressive — in those contexts almost nothing invisible belongs, and the cost of a stray character is high.
Frequently asked questions
Do hidden characters mean the text was written by AI?
No. This is a widespread misconception. Language models produce ordinary text and do not insert zero-width markers as a signature. These characters come from PDFs, word processors, web pages, and encoding conversions.
Is it safe to remove every invisible character?
Not always. Zero-width joiners and non-joiners are required for correct rendering in Arabic, Persian, and Indic scripts, and for building emoji sequences. Directional controls keep mixed-direction text readable. Remove those only when you can verify they are wrong.
Why does my text look identical but compare as different?
Usually a non-breaking space where an ordinary space belongs, or a zero-width character inside a word. Both render identically and compare as unequal, which is exactly why an inspection tool is needed.
Is my pasted text sent to a server?
No. The scanner is a pure function running in your browser. The text stays in the page's memory, is never transmitted, and is discarded when you close or reload the tab.
What is a byte order mark and why is it in my text?
U+FEFF signals byte order in some encodings. It commonly survives file conversion and ends up at the start of pasted text, where it can break parsers and exact matching. It is safe to remove from running text.
Primary sources
The technical claims on this page follow the published specifications below rather than our own assertions.
- Unicode Character Database (UAX #44)
Defines the general categories, including the Cf format characters this scanner names.
- Unicode Bidirectional Algorithm (UAX #9)
Explains embedding, override, and isolate characters, and why removing them changes how text reads.
- Unicode Security Mechanisms (UTS #39)
Covers confusable and invisible characters as a spoofing and review-evasion concern.