UTF-8 Encoding Table

–ranges and examples
FromToBytesByte 1Byte 2Byte 3Byte 4Covers
U+000000U+00007F10''yyyzzzz''128 code points - ASCII, byte-for-byte identical
U+000080U+0007FF2110''xxxyy''10''yyzzzz''1,920 code points - Latin, Greek, Cyrillic, Hebrew, Arabic, Armenian and neighbors
U+000800U+00FFFF31110''wwww''10''xxxxyy''10''yyzzzz''61,440 code points - the rest of the BMP, most CJK
U+010000U+10FFFF411110''uvv''10''vvwwww''10''xxxxyy''10''yyzzzz''1,048,576 code points - emoji, rare CJK, everything beyond the BMP
CharCode pointBytesUTF-8 hex
$U+0024124
éU+00E92C3 A9
€U+20AC3E2 82 AC
—U+20143E2 80 94
’U+20193E2 80 99
中U+4E2D3E4 B8 AD
😀U+1F6004F0 9F 98 80
桁U+68413E6 A1 81
UTF-8 is the encoding of the web: as of 2026, Wikipedia puts 99.1% of webpages on it (versus the 9.3% legacy holdout of Windows-1252 charted on this site). Its design is ASCII backwards compatibility: “the first 128 characters of Unicode... are encoded using a single byte with the same binary value as ASCII, so that a UTF-8-encoded file using only those characters is identical to an ASCII file.” Above that, the lead-byte pattern is the whole grammar: 110/1110/11110 prefixes announce 2-, 3- and 4-byte units, and every continuation byte starts with 10 - which makes UTF-8 a prefix code, self-synchronizing (a decoder can find the start of any character from a random position by backing up at most three bytes) and order-preserving (sorting UTF-8 strings sorts them by code point).
Wikipedia’s worked example is reproduced exactly: the character 栙 (U+6841) is 0110 1000 0100 0001 in binary, which makes its UTF-8 encoding 11100110 10100001 10000001 - hex E6 A1 81, the last row of the example table (every example row here is byte-verified, not transcribed). The capacity rows are the standard’s own arithmetic: 128 one-byte ASCII slots, 1,920 two-byte slots (Latin/Greek/Cyrillic/Hebrew/Arabic/Armenian and neighbors), 61,440 three-byte slots (the rest of the BMP, most CJK), and 1,048,576 four-byte slots (emoji and everything beyond the BMP). Adoption note, verbatim from Wikipedia: C++23 adopted UTF-8 as the only portable source-and-execution character set, Python 3.15 makes it the I/O default, and since May 2019 Windows lets applications set UTF-8 as their code page - with Microsoft itself calling UTF-16 “a unique burden that Windows places on code that targets multiple platforms.” Reference: the Unicode Standard, latest version. Sibling charts: the Windows-1252 table (the legacy byte it replaced) and the JS encoding table (the TextEncoder API that produces these bytes in the browser).

UTF-8 is the default encoding of the digital world - Wikipedia’s September 2026 count puts it on 99.1% of webpages - and its entire design fits in one table: code points U+0000–U+10FFFF, encoded in one to four bytes, with the lead byte announcing the length and every continuation byte starting with binary 10.

The genius constraint is history: UTF-8 was designed for backward compatibility with ASCII, so “the first 128 characters of Unicode... are encoded using a single byte with the same binary value as ASCII,” making pure-ASCII files byte-identical in both encodings. Every byte above 0x7F belongs to a multi-byte sequence - which is why a stray 0x93 from a Windows-1252 file is invalid UTF-8, and why the two encodings charted on this site interlock so cleanly.

Below: the four range rows of the byte grammar, then a worked example table - dollar sign to emoji to Wikipedia’s own 栙 (U+6841) - each row byte-verified rather than transcribed, so the hex you see is what any TextEncoder produces.

How to use

  1. Compute a byte count in one glance: find the code point’s range row and the Bytes column is your answer - U+0041 needs 1 byte, U+00E9 needs 2, U+4E2D needs 3, U+1F600 needs 4. The lead byte pattern (0..., 110..., 1110..., 11110...) is redundant with the range, which is what makes broken sequences detectable.
  2. Decode mojibake by re-encoding: if text displays as &#195;&#169; where &#233; should be, the bytes are UTF-8 (C3 A9) being read as Latin-1/Windows-1252; the <a href="https://tooldune.com/windows-1252-table/" rel="noopener">Windows-1252 table</a> on this site lists that range byte by byte for the reverse trip.
  3. Validate hand-assembled sequences against the grammar: lead byte 110xxxx must be followed by exactly one 10xxxxxx byte, 1110 by two, 11110 by three - a sequence like E0 80 80 (three bytes encoding U+0000) is invalid because the code point must use the shortest form, and continuation bytes never appear without a lead.

Frequently asked questions

How many bytes does a character use in UTF-8?

One to four, by range: U+0000&#8211;U+007F is one byte (ASCII, identical bytes), U+0080&#8211;U+07FF is two, U+0800&#8211;U+FFFF is three, and U+10000&#8211;U+10FFFF is four. The buckets hold 128, 1,920, 61,440 and 1,048,576 code points respectively - Wikipedia&#8217;s own arithmetic, reproduced in the first table.

Why is UTF-8 called UTF-8?

Unicode Transformation Format, 8-bit: the name comes from the standard it encodes and its 8-bit code units. It replaced the failed UTF-1 (Wikipedia&#8217;s history section: the ISO 10646 draft&#8217;s encoding &#8220;was not satisfactory on performance grounds&#8221; and lacked a clean ASCII boundary), and UTF-8&#8217;s self-synchronizing design fixed both problems.

What makes UTF-8 self-synchronizing?

The bit patterns do the bookkeeping: lead bytes start 0, 110, 1110 or 11110 and continuations always start 10, so a decoder can back up at most three bytes from any random position to find a character boundary - impossible in Shift-JIS-era encodings. Wikipedia also notes the design keeps UTF-8 string sorting identical to code-point sorting.

Is UTF-8 the same as Unicode?

No - Unicode is the character set (the code points), UTF-8 is one of its encodings (the bytes). UTF-16 and UTF-32 encode the same code points differently; UTF-8 won the web (99.1% of pages per Wikipedia, 2026) because of its ASCII compatibility and byte economy for Latin text, and because C++23 adopted it as the only portable source encoding while Python 3.15 makes it the default I/O encoding.

Why does &#8364; take three bytes but Windows-1252 does it in one?

Because the encodings solve different problems: Windows-1252 spends a scarce single-byte slot (0x80) on it, while UTF-8&#8217;s U+20AC lives in the three-byte range (E2 82 AC). The single-byte version only covers 256 characters total - the moment your text needs both &#8364; and &#20013;, one byte per character stops being an option, and the <a href="https://tooldune.com/windows-1252-table/" rel="noopener">Windows-1252 table</a> shows exactly which 32 slots it borrowed to get as far as it did.

Related tools