JavaScript Intl.Segmenter Table
| Piece | What it does | Field note |
|---|---|---|
Intl.Segmenter(lo, {granularity}) | Locale-aware splitting | grapheme / word / sentence - three granularities |
grapheme mode | User-visible characters | ๐จโ๐ฉโ๐ง = ONE grapheme, 8 code points - split() lies, Segmenter does not |
segment.segment(str) | The iteration | Segment objects with .segment/.index - or spread for arrays |
word mode + isWordLike | Real words vs spaces | Filters punctuation - CJK finally has word boundaries |
sentence mode | Locale sentence rules | Abbreviations and CJK ๅฅๅท handled - hard to hand-roll |
Loose vs strict | Legacy safety | strict keeps old (wrong) behavior for compat - use loose |
vs str.split('') | Code points vs people | split breaks emoji families and combining marks apart |
cursor/selection math | The editor use | Left/right arrow moves by GRAPHEME - text editing's ground truth |
Intl.Segmenter splits text the way USERS see it: grapheme mode treats an emoji family (๐จโ๐ฉโ๐ง - eight code points) as ONE character where str.split('') shreds it into replacement-marked fragments; word mode finds real words including CJK; sentence mode applies locale sentence rules.
Bottom line: cursor movement, character counters and truncation must count GRAPHEMES, not code points - the text-editor left-arrow truth is that the user's 'character' is a grapheme. And isWordLike filters the punctuation segments that word mode also emits.
The honest part: Segmenter is locale-aware by construction - the same string segments differently under different locales (Japanese word boundaries differ from German), which is the point: it is the Intl family, and the segmentation follows the locale's linguistic data, not a universal rule.
How to use
- Count what users count: [...new Intl.Segmenter('en', {granularity:'grapheme'}).segment(str)].length - the emoji-safe character counter in one line.
- Truncate without shredding: collect grapheme segments and join the first N - '๐จโ๐ฉโ๐ง...'.slice(0, 8) produces broken output; segments never do.
- Extract real words: [...seg.segment(text)].filter(s => s.isWordLike).map(s => s.segment) - CJK text finally yields word lists for tagging and search.
Frequently asked questions
Why does splitting an emoji break my string, and how does Segmenter fix it?
A rendered 'character' is a GRAPHEME - one or more code points composed: an emoji family is eight code points (three people joined by zero-width joiners), and accented letters can be two (letter plus combining mark). split('') and code-point iteration cut between JOINERS and combining marks, producing lone surrogates and naked combining marks that render as replacement boxes. Intl.Segmenter with granularity 'grapheme' walks the Unicode grapheme cluster rules: the family comes out as one segment, combining marks stay attached, and a character counter built on segments reports what users would count with their eyes. This is the fix for every truncated-string-shows-boxes bug.
How does word segmentation handle CJK and what is isWordLike?
Word boundaries need language data. European text can approximate words with spaces, but Chinese and Japanese have NO spaces - only a locale's linguistic data knows where ่ฏ end. Intl.Segmenter's word granularity uses that data: Japanese text yields real word segments where split(' ') returns the whole sentence. The nuance: word mode emits segments for punctuation and spaces TOO, each flagged isWordLike: false - so real-word extraction is the one-line filter [...seg.segment(text)].filter(s => s.isWordLike). This is the difference between a search index that matches ่ฏๆฑ and one that matches whole sentences.
What do loose and strict mean for legacy compatibility?
They control whether Segmenter may CORRECT old behavior. Some engines' earlier (non-standard) segmentation implementations had quirks that shipped software came to depend on; granularity 'strict' preserves the old (wrong by Unicode) behavior for that compatibility, while 'loose' - the default recommendation - follows the current Unicode rules. For new code the answer is always loose: strict exists so a engine update cannot silently change an old program's segmentation mid-flight. The general lesson repeats across the Intl family: linguistic data and rules evolve, and the API surfaces escape hatches rather than breaking deployed behavior.
Where does cursor movement meet segmentation in real products?
Every text input. Pressing left-arrow should move one USER-visible character: over an emoji family, over 'e' plus its combining accent, over a Devanagari cluster - which is exactly grapheme segmentation, and browsers use these cluster rules natively for their own inputs. The moment YOU render and edit text yourself (canvas editors, custom rich-text boxes, terminal emulators), the responsibility transfers: cursor columns computed in code-point offsets skip into the middle of emoji, deleting one 'character' via code points deletes a joint. Segmenter is the ground truth library for converting 'one keypress' into 'one code-point range' - the reason every editor team lists it as a dependency or an early primitive.