UTF-8 Encoding Table
| From | To | Bytes | Byte 1 | Byte 2 | Byte 3 | Byte 4 | Covers |
|---|---|---|---|---|---|---|---|
| U+000000 | U+00007F | 1 | 0''yyyzzzz'' | 128 code points - ASCII, byte-for-byte identical | |||
| U+000080 | U+0007FF | 2 | 110''xxxyy'' | 10''yyzzzz'' | 1,920 code points - Latin, Greek, Cyrillic, Hebrew, Arabic, Armenian and neighbors | ||
| U+000800 | U+00FFFF | 3 | 1110''wwww'' | 10''xxxxyy'' | 10''yyzzzz'' | 61,440 code points - the rest of the BMP, most CJK | |
| U+010000 | U+10FFFF | 4 | 11110''uvv'' | 10''vvwwww'' | 10''xxxxyy'' | 10''yyzzzz'' | 1,048,576 code points - emoji, rare CJK, everything beyond the BMP |
| Char | Code point | Bytes | UTF-8 hex |
|---|---|---|---|
| $ | U+0024 | 1 | 24 |
| é | U+00E9 | 2 | C3 A9 |
| € | U+20AC | 3 | E2 82 AC |
| — | U+2014 | 3 | E2 80 94 |
| ’ | U+2019 | 3 | E2 80 99 |
| 中 | U+4E2D | 3 | E4 B8 AD |
| 😀 | U+1F600 | 4 | F0 9F 98 80 |
| 桁 | U+6841 | 3 | E6 A1 81 |
UTF-8 is the default encoding of the digital world - Wikipedia’s September 2026 count puts it on 99.1% of webpages - and its entire design fits in one table: code points U+0000–U+10FFFF, encoded in one to four bytes, with the lead byte announcing the length and every continuation byte starting with binary 10.
The genius constraint is history: UTF-8 was designed for backward compatibility with ASCII, so “the first 128 characters of Unicode... are encoded using a single byte with the same binary value as ASCII,” making pure-ASCII files byte-identical in both encodings. Every byte above 0x7F belongs to a multi-byte sequence - which is why a stray 0x93 from a Windows-1252 file is invalid UTF-8, and why the two encodings charted on this site interlock so cleanly.
Below: the four range rows of the byte grammar, then a worked example table - dollar sign to emoji to Wikipedia’s own 栙 (U+6841) - each row byte-verified rather than transcribed, so the hex you see is what any TextEncoder produces.
How to use
- Compute a byte count in one glance: find the code point’s range row and the Bytes column is your answer - U+0041 needs 1 byte, U+00E9 needs 2, U+4E2D needs 3, U+1F600 needs 4. The lead byte pattern (0..., 110..., 1110..., 11110...) is redundant with the range, which is what makes broken sequences detectable.
- Decode mojibake by re-encoding: if text displays as é where é should be, the bytes are UTF-8 (C3 A9) being read as Latin-1/Windows-1252; the <a href="https://tooldune.com/windows-1252-table/" rel="noopener">Windows-1252 table</a> on this site lists that range byte by byte for the reverse trip.
- Validate hand-assembled sequences against the grammar: lead byte 110xxxx must be followed by exactly one 10xxxxxx byte, 1110 by two, 11110 by three - a sequence like E0 80 80 (three bytes encoding U+0000) is invalid because the code point must use the shortest form, and continuation bytes never appear without a lead.
Frequently asked questions
How many bytes does a character use in UTF-8?
One to four, by range: U+0000–U+007F is one byte (ASCII, identical bytes), U+0080–U+07FF is two, U+0800–U+FFFF is three, and U+10000–U+10FFFF is four. The buckets hold 128, 1,920, 61,440 and 1,048,576 code points respectively - Wikipedia’s own arithmetic, reproduced in the first table.
Why is UTF-8 called UTF-8?
Unicode Transformation Format, 8-bit: the name comes from the standard it encodes and its 8-bit code units. It replaced the failed UTF-1 (Wikipedia’s history section: the ISO 10646 draft’s encoding “was not satisfactory on performance grounds” and lacked a clean ASCII boundary), and UTF-8’s self-synchronizing design fixed both problems.
What makes UTF-8 self-synchronizing?
The bit patterns do the bookkeeping: lead bytes start 0, 110, 1110 or 11110 and continuations always start 10, so a decoder can back up at most three bytes from any random position to find a character boundary - impossible in Shift-JIS-era encodings. Wikipedia also notes the design keeps UTF-8 string sorting identical to code-point sorting.
Is UTF-8 the same as Unicode?
No - Unicode is the character set (the code points), UTF-8 is one of its encodings (the bytes). UTF-16 and UTF-32 encode the same code points differently; UTF-8 won the web (99.1% of pages per Wikipedia, 2026) because of its ASCII compatibility and byte economy for Latin text, and because C++23 adopted it as the only portable source encoding while Python 3.15 makes it the default I/O encoding.
Why does € take three bytes but Windows-1252 does it in one?
Because the encodings solve different problems: Windows-1252 spends a scarce single-byte slot (0x80) on it, while UTF-8’s U+20AC lives in the three-byte range (E2 82 AC). The single-byte version only covers 256 characters total - the moment your text needs both € and 中, one byte per character stops being an option, and the <a href="https://tooldune.com/windows-1252-table/" rel="noopener">Windows-1252 table</a> shows exactly which 32 slots it borrowed to get as far as it did.