JavaScript Encoding Table
| Piece | Direction | Field note |
|---|---|---|
new TextEncoder().encode(s) | String to UTF-8 bytes | Returns a Uint8Array; the instance is stateless and reusable forever |
new TextDecoder().decode(bytes) | Bytes to string | UTF-8 default; label param decodes legacy charsets like windows-1252 |
decode(bytes, {stream: true}) | Chunked input | Carries split multi-byte tails to the next call - chunk borders no longer corrupt emoji |
{fatal: true} | Strict decode | Bad bytes THROW instead of planting invisible U+FFFD replacement characters |
btoa / atob | Base64, Latin-1 ONLY | Any char above U+00FF makes btoa throw InvalidCharacterError - every emoji, every CJK char |
TextEncoder then base64 | The safe two-step | String to UTF-8 bytes FIRST, then base64 the bytes - chunked fromCharCode for big strings |
bytes.byteLength vs s.length | Bytes vs code units | UTF-8 bytes differ from UTF-16 units on every non-Latin char - upload math lives in bytes |
base64url variant | URL-safe form | Swap + to -, / to _, drop padding - or tokens break in query strings (+ reads as space) |
TextEncoder and TextDecoder are the browser's canonical UTF-8 bridge: encoder.encode('hi') turns a string into a Uint8Array of bytes, decoder.decode(bytes) turns bytes back - stateless, reusable, and the ONLY pair you need for text-to-bytes in modern code. FileReader async dances and hand-rolled UTF-8 math are legacy.
Bottom line: atob and btoa are NOT string-to-base64 - they are Latin-1-to-base64, and any character above U+00FF (every emoji, every CJK char) makes btoa THROW. The safe path is two steps: TextEncoder to bytes, then base64 the bytes (btoa(String.fromCharCode(...new Uint8Array(bytes))) for small strings, chunked for big ones) - or simply let a library do it and stop seeing InvalidCharacterError in production.
The honest part: decode errors are silent by default - TextDecoder swaps bad bytes for U+FFFD replacement characters unless you pass { fatal: true }, which turns corruption into an exception you can actually catch. For network chunks, decode(bytes, { stream: true }) carries partial multi-byte sequences across calls; without it, a 3-byte emoji split across two chunks decodes as two garbage halves.
How to use
- Encode to bytes: const bytes = new TextEncoder().encode('hello') - UTF-8 always, one reusable instance, no options worth setting (UTF-16 LE via TextEncoder is not a thing; use typed arrays manually if you truly need it).
- Decode strictly: new TextDecoder('utf-8', { fatal: true }).decode(bytes) - throws on malformed input instead of planting invisible U+FFFD characters that corrupt downstream string comparison.
- Decode network chunks: decoder.decode(chunk, { stream: true }) per chunk, one final decode() without the flag at the end - multi-byte characters split across chunks survive intact.
- Base64 safely: btoa(String.fromCharCode(...bytes)) for short strings; for anything user-sized, loop 0x8000 chars per fromCharCode batch or reach for Uint8Array.toBase64() on fresh engines.
- Know your sizes: bytes.byteLength counts UTF-8 bytes; string.length counts UTF-16 units - upload limits and Content-Length live in the byte world, and the two diverge on every emoji and non-Latin script.
Frequently asked questions
Why does btoa throw InvalidCharacterError on emoji but atob works fine?
btoa accepts only Latin-1: characters U+0000 to U+00FF. An emoji is code points far above that, so the function refuses - it has no encoding step to fall back on. atob rarely throws because its OUTPUT range (Latin-1) is what it produces. The fix is never try/catch around btoa - it is encoding the string to UTF-8 bytes first (TextEncoder), then base64-encoding the BYTES; the error is a category error, not a runtime fluke, and chunked fromCharCode keeps large inputs from hitting the spread-argument limit.
When do I need the stream option on TextDecoder?
Whenever bytes arrive in pieces: fetch body readers, WebSockets with binary frames, file slices. UTF-8 characters are 1-4 bytes and do not respect chunk boundaries - a chunk may end mid-character. decode(chunk, { stream: true }) buffers the incomplete tail and prepends it to the next call, so the full stream decodes byte-exact. Without it, every boundary-split character becomes replacement garbage; the same option exists on TextEncoder for the encode direction.
Is atob/btoa output URL-safe?
No - standard base64 carries + and / and = padding, all of which mean things in URLs and filenames. URL-safe base64 (base64url) swaps - for + and _ for / and drops padding; you must do that replace yourself, or use the engines that ship base64url natively (JWT libraries wrap this for a reason). Common bug: base64ing a token into a query string unescaped, where + decodes back as a space and validation fails server-side.
What encoding should I actually use in 2026?
UTF-8, always, everywhere: TextEncoder only produces UTF-8, fetch and FormData assume it, databases default to it, and the Web platform standardized on it years ago. The remaining legacy spaces are Windows internal APIs (UTF-16) and protocol headers (Latin-1), which is exactly what TextDecoder's label parameter is for - new TextDecoder('windows-1252') decodes legacy bytes when ingesting old files. Never build new output in a legacy encoding; decode legacy input on the way in and let UTF-8 be the only encoding your system speaks.