
Who it is for
It is for editors, teachers, developers and anyone security-minded who has been handed text and wants to know exactly what is in it. It is not an AI detector, and it says so: it reports characters, offsets and counts, and gives no score or probability about who or what wrote the words.
How it works
Everything runs in the tab. The script’s header says it contains no fetch, no XMLHttpRequest, no beacon, no WebSocket and no form post; the only file access is a FileReader on a file you choose. We searched the served app.js (122,276 bytes) and found none of those calls.
One pass over the text’s code points feeds eight detectors: invisible and format characters; three kinds of payload (Unicode tag characters, variation selectors and zero-width bit channels); homoglyphs and mixed scripts; bidirectional control abuse; a typographic fingerprint; whitespace; normalisation against NFC and NFKC; and text hidden in HTML. Tag characters in the range U+E0020–E007F are decoded back into the ASCII they smuggle.
Safe mode is on by default. It keeps characters that have an ordinary job — emoji sequences, subdivision flags, Indic and Arabic shaping, CJK variation selectors — so cleaning a passage does not break it. And the counting carries on after the listing stops: the display caps at 200 rows, but 5,000 findings are reported as 5,000.
What we tested
All times are UTC on 27 September 2026, measured against a copy of labs.llc build 537 served from our own machine, or against the public source the page itself calls.
- Page and script, searched for network callsPage 200, 74,229 bytes; app.js 200, 122,276 bytes; no fetch(, XMLHttpRequest( or sendBeacon( anywhere.
- The carrier pattern on a crafted 26-code-point stringSix hits: U+200B, U+00AD, U+202E, U+202C, U+E0068, U+E0069. The Cyrillic “а” in “аpple” was left to the homoglyph detector, as designed.
- The lookalike table, counted57 pairs across Cyrillic, Greek, Armenian and Cherokee.
The crafted string mixed a zero-width space, a soft hyphen, a right-to-left override and its closing mark, two tag characters spelling “hi”, and a Cyrillic letter inside an English word. We ran the stripper’s pattern from line 31 of app.js in JavaScriptCore.
How it compares
We checked each alternative’s own pages on 27 September 2026 and describe only what we saw there. Each gets the pick below when it does the job better.
ASCII Smuggler (Embrace The Red)
checked 27 Sep 2026A browser tool that turns text into invisible Unicode and decodes it back. It handles Unicode tags, variant selectors, “sneaky bits” and zero-width characters, with highlight and auto-decode modes and a debug view.
Better at
- It encodes as well as decodes, so you can build a test payload from your own message.
- A debug view of the encoding in binary, hex or Unicode.
Where Watermark goes further
- Seven detectors beyond payloads, each switchable and each reporting offsets.
- Safe mode that protects emoji, flags and shaping, plus a marked diff of what was removed.
- JSON and Markdown reports, and a bookmarklet that strips the page you are reading.
invisible-characters.com
checked 27 Sep 2026A reference to more than 80 invisible Unicode characters, with a decoder that shows which of them a block of text contains. It reveals them; it does not remove them.
Better at
- A browsable catalogue with explanations — a good place to learn what each character is.
Where Watermark goes further
- Removes what it finds, with safe mode and a diff, and reports offsets.
- Catches visible deceptions too: lookalikes, bidi reordering, normalisation and hidden HTML.
r12a Unicode code converter
checked 27 Sep 2026A converter between Unicode representations: code points, UTF-8 and UTF-16, JavaScript, CSS, Perl and Rust escapes, HTML references and percent-encoding, with options to escape invisible characters and bidi controls.
Better at
- Exact conversion into a dozen escape and encoding forms, which Watermark does not offer.
- Reverse conversion from bare hex or decimal numbers.
Where Watermark goes further
- Groups findings by kind with counts and offsets, and says which characters have innocent jobs, instead of escaping everything.
Where it falls short
- The lookalike table is 57 hand-picked pairs from four scripts, not Unicode’s full confusables data, so a lookalike from any other script passes.
- Only mixed-script words are flagged. A word spelled entirely in Cyrillic lookalikes gets past the homoglyph detector — deliberately, so that Russian writers are not accused of an attack, but it is still a gap.
- There is no encoder for your own message. Its samples carry fixed strings such as “LABS-2026”; ASCII Smuggler will encode anything you type.
- It cannot say whether a person or a model wrote the text. Anyone arriving for that answer leaves without it.
Which one to pick
| If you need… | Pick |
|---|---|
| Find and remove invisible characters, with positions | Watermarkours |
| Inspect text without it leaving your machine | Watermarkours |
| Make a hidden-text payload to test a system | ASCII Smuggler |
| Look up what one invisible character is for | invisible-characters.com |
| Turn text into escapes or UTF-8 bytes | r12a converter |
The fit
Best for
- Cleaning copied text before it goes into a document, a form or code
- Spotting smuggled tag-character or zero-width payloads — and reading what they say
- Editors who want offsets, not a verdict
Not for
- Deciding whether an AI wrote something
- Producing encoded payloads of your own
- Catching every possible lookalike in every script
Pick Watermark if you need to know exactly which invisible characters are in a passage, and where they sit.
Sources for this review
- watermark/assets/js/app.js in build 537 — header lines 3–24, CORE_RE at line 31, nameOf() at 80–83, caps at 310–314, homoglyph rule at 759–825, detector table at 1381–1388
- ASCII Smuggler, invisible-characters.com and the r12a converter, each fetched on 27 September 2026 at about 20:03 UTC