Unicode Character Inspector

–characters, bytes
Invisible suspectCodepointUTF-8Where it hides
Zero-width spaceU+200BE2 80 8Bcopied web text, fake spaces
Non-breaking spaceU+00A0C2 A0PDF and CMS pastes
Zero-width joinerU+200DE2 80 8Demoji sequences, homoglyph tricks
Byte-order markU+FEFFEF BB BFfile starts, database exports
Soft hyphenU+00ADC2 ADword processors, scraped titles
Byte lengths per RFC 3629 (UTF-8): 1 byte to U+007F, 2 to U+07FF, 3 to U+FFFF, 4 above (emoji included). Superscripts and math symbols: the ASCII table, text to binary.

Every character you type is a number underneath. This inspector takes your text apart character by character and shows each one's Unicode codepoint (U+0041-style), its decimal value, and the exact byte sequence UTF-8 uses to store it - the trio that explains nearly every 'why does this string behave weird' mystery. Paste suspicious text from a form field, a copied PDF, or a database export and the invisible characters - zero-width spaces, non-breaking spaces, directional marks - light up in the table next to the letters you can see.

The encoding column follows RFC 3629, the specification that fixed UTF-8: characters up to U+007F take one byte, U+0080-U+07FF take two, U+0800-U+FFFF take three, and everything above (including emoji, which live past U+FFFF in the astral planes) takes four. That is why an emoji costs four bytes in a database column sized for three - the inspector makes the arithmetic visible one character at a time.

How to use

  1. Paste or type text in the box - the table below lists each character with its display, codepoint, decimal value and UTF-8 bytes.
  2. Look for empty-looking rows: U+200B (zero-width space), U+00A0 (non-breaking space) and U+FEFF (byte-order mark) are the usual invisible suspects.
  3. Watch the byte count climb past U+FFFF: emoji and rare CJK characters take four UTF-8 bytes, which is what breaks 3-byte-sized database columns.

Frequently asked questions

How do I find invisible characters in text?

Paste the text here and scan the codepoint column: visible letters and punctuation have familiar values, while the invisible classics have none - U+200B zero-width space, U+200C and U+200D joiners, U+00A0 non-breaking space and U+FEFF byte-order mark all render as blank or empty in the display column but show real codepoints and byte counts in the table. Any row where the display looks empty but bytes are nonzero is your culprit.

Why does one emoji show as two characters?

Emoji beyond U+FFFF are stored as surrogate pairs in JavaScript-style UTF-16: two 16-bit halves that this inspector joins back into one character using codePointAt, which is why an emoji occupies one row, one codepoint - and four UTF-8 bytes. If another tool counts your emoji as two characters, it is counting surrogates; databases and JSON count codepoints, which is the number this tool shows.

What is the difference between a codepoint and a byte?

A codepoint is the abstract number Unicode assigns a character (U+0041 for A); bytes are how that number is physically stored. UTF-8 stores A as one byte 0x41, รฉ as two bytes 0xC3 0xA9, and the emoji ๐ŸŽ‰ as four bytes 0xF0 0x9F 0x8E 0x89. The inspector shows both columns so you can see that 'same characters, different bytes' - the usual cause of 'identical' strings failing to match - is an encoding difference, not a spelling one.

Can this inspect an entire file?

For pasted text, yes - the table covers what fits in the box, which handles names, form values and error strings, the places invisible characters actually strike. Whole-file analysis is a different job: run a hex editor or an encoding linter over the bytes. The pasted-snippet approach matches how these bugs surface anyway - one weird character in one field, not a whole document of them.

Related tools