JavaScript Encoding Table

PieceDirectionField note
new TextEncoder().encode(s)String to UTF-8 bytesReturns a Uint8Array; the instance is stateless and reusable forever
new TextDecoder().decode(bytes)Bytes to stringUTF-8 default; label param decodes legacy charsets like windows-1252
decode(bytes, {stream: true})Chunked inputCarries split multi-byte tails to the next call - chunk borders no longer corrupt emoji
{fatal: true}Strict decodeBad bytes THROW instead of planting invisible U+FFFD replacement characters
btoa / atobBase64, Latin-1 ONLYAny char above U+00FF makes btoa throw InvalidCharacterError - every emoji, every CJK char
TextEncoder then base64The safe two-stepString to UTF-8 bytes FIRST, then base64 the bytes - chunked fromCharCode for big strings
bytes.byteLength vs s.lengthBytes vs code unitsUTF-8 bytes differ from UTF-16 units on every non-Latin char - upload math lives in bytes
base64url variantURL-safe formSwap + to -, / to _, drop padding - or tokens break in query strings (+ reads as space)
Reference: the MDN Encoding API guide. TextEncoder and TextDecoder are the canonical UTF-8 bridge - stateless, reusable, and all modern code needs. The classic production crash is btoa on user input: atob and btoa predate the Encoding API and speak Latin-1 only, so the fix is architectural, not try/catch - encode to UTF-8 bytes first, base64 the bytes second.
Bottom line: decode errors are SILENT by default (U+FFFD swaps) - pass fatal: true whenever corrupt bytes should be a caught exception instead of invisible garbage that breaks string comparison downstream. And the stream option is not optional for network paths: UTF-8 characters are 1-4 bytes and ignore chunk boundaries, so unflagged chunked decoding shreds boundary-straddling characters into replacement halves.
Related tools: the typed arrays table (the Uint8Array the encoder returns into), the streams table (where chunked decode actually happens), the compression streams table (gzip after encode), the Blob table (bytes with a MIME type), and the File API table (reading files into the byte world).

TextEncoder and TextDecoder are the browser's canonical UTF-8 bridge: encoder.encode('hi') turns a string into a Uint8Array of bytes, decoder.decode(bytes) turns bytes back - stateless, reusable, and the ONLY pair you need for text-to-bytes in modern code. FileReader async dances and hand-rolled UTF-8 math are legacy.

Bottom line: atob and btoa are NOT string-to-base64 - they are Latin-1-to-base64, and any character above U+00FF (every emoji, every CJK char) makes btoa THROW. The safe path is two steps: TextEncoder to bytes, then base64 the bytes (btoa(String.fromCharCode(...new Uint8Array(bytes))) for small strings, chunked for big ones) - or simply let a library do it and stop seeing InvalidCharacterError in production.

The honest part: decode errors are silent by default - TextDecoder swaps bad bytes for U+FFFD replacement characters unless you pass { fatal: true }, which turns corruption into an exception you can actually catch. For network chunks, decode(bytes, { stream: true }) carries partial multi-byte sequences across calls; without it, a 3-byte emoji split across two chunks decodes as two garbage halves.

How to use

  1. Encode to bytes: const bytes = new TextEncoder().encode('hello') - UTF-8 always, one reusable instance, no options worth setting (UTF-16 LE via TextEncoder is not a thing; use typed arrays manually if you truly need it).
  2. Decode strictly: new TextDecoder('utf-8', { fatal: true }).decode(bytes) - throws on malformed input instead of planting invisible U+FFFD characters that corrupt downstream string comparison.
  3. Decode network chunks: decoder.decode(chunk, { stream: true }) per chunk, one final decode() without the flag at the end - multi-byte characters split across chunks survive intact.
  4. Base64 safely: btoa(String.fromCharCode(...bytes)) for short strings; for anything user-sized, loop 0x8000 chars per fromCharCode batch or reach for Uint8Array.toBase64() on fresh engines.
  5. Know your sizes: bytes.byteLength counts UTF-8 bytes; string.length counts UTF-16 units - upload limits and Content-Length live in the byte world, and the two diverge on every emoji and non-Latin script.

Frequently asked questions

Why does btoa throw InvalidCharacterError on emoji but atob works fine?

btoa accepts only Latin-1: characters U+0000 to U+00FF. An emoji is code points far above that, so the function refuses - it has no encoding step to fall back on. atob rarely throws because its OUTPUT range (Latin-1) is what it produces. The fix is never try/catch around btoa - it is encoding the string to UTF-8 bytes first (TextEncoder), then base64-encoding the BYTES; the error is a category error, not a runtime fluke, and chunked fromCharCode keeps large inputs from hitting the spread-argument limit.

When do I need the stream option on TextDecoder?

Whenever bytes arrive in pieces: fetch body readers, WebSockets with binary frames, file slices. UTF-8 characters are 1-4 bytes and do not respect chunk boundaries - a chunk may end mid-character. decode(chunk, { stream: true }) buffers the incomplete tail and prepends it to the next call, so the full stream decodes byte-exact. Without it, every boundary-split character becomes replacement garbage; the same option exists on TextEncoder for the encode direction.

Is atob/btoa output URL-safe?

No - standard base64 carries + and / and = padding, all of which mean things in URLs and filenames. URL-safe base64 (base64url) swaps - for + and _ for / and drops padding; you must do that replace yourself, or use the engines that ship base64url natively (JWT libraries wrap this for a reason). Common bug: base64ing a token into a query string unescaped, where + decodes back as a space and validation fails server-side.

What encoding should I actually use in 2026?

UTF-8, always, everywhere: TextEncoder only produces UTF-8, fetch and FormData assume it, databases default to it, and the Web platform standardized on it years ago. The remaining legacy spaces are Windows internal APIs (UTF-16) and protocol headers (Latin-1), which is exactly what TextDecoder's label parameter is for - new TextDecoder('windows-1252') decodes legacy bytes when ingesting old files. Never build new output in a legacy encoding; decode legacy input on the way in and let UTF-8 be the only encoding your system speaks.

Related tools