What are zero-width characters?
Unicode includes dozens of invisible characters that render as nothing but still exist in your data: zero-width space (U+200B), zero-width non-joiner (U+200C), zero-width joiner (U+200D), word joiner (U+2060), soft hyphen (U+00AD), BOM (U+FEFF), left-to-right / right-to-left marks (U+200E / U+200F), and more. They're used legitimately in text shaping — and used maliciously to fingerprint text, break filters, or smuggle payloads past input validation.
Where they hide
- LLM output occasionally includes U+200B or soft hyphens that break regex matches and exact string comparisons.
- Scraped pages often carry BOMs at the start of files and zero-width chars from CMS templates.
- Copied code can contain invisible chars that break identifiers or JSON parsing with "unexpected character" errors.
- Spam and phishing use zero-width spaces to defeat keyword filters ("freе" with a homoglyph + ZWJ).
- Watermarking — embedding hidden characters in text to trace copies (the scanner will surface these).
How to use the results
Scan highlights every invisible character with a visible marker and lists each code point with its name and count. Remove all strips them — but be careful: if the source used them as watermarks, you're removing exactly what someone planted. For JSON, code or DB keys, removing is usually safe; for prose, check the soft hyphens didn't carry intended line-break hints.