Unicode Categories Table
| Code | Category | Examples | What belongs here | Assigned |
|---|---|---|---|---|
| Lu | Uppercase Letter | A É Θ | Capital letters across all alphabets | 1,906 |
| Ll | Lowercase Letter | a é θ | Small letters across all alphabets | 2,366 |
| Lt | Titlecase Letter | Dž Lj | Digraph letters like Dz that title-case half | 31 |
| Lm | Modifier Letter | ˆ ʰ | Letters that modify the next sound, not standalone words | 473 |
| Lo | Other Letter | 中 あ א | Every ideograph, kana, abjad letter - the biggest bucket | 21,852 |
| Mn | Nonspacing Mark | ́ ̈ | Accents that overlay the previous character | 2,090 |
| Mc | Spacing Mark | ः ั | Combining marks that claim their own width | 477 |
| Me | Enclosing Mark | ⃝ | Rings and circles that wrap the previous character | 13 |
| Nd | Decimal Number | 0 3 २ | Digits usable in decimal numbers, any script | 770 |
| Nl | Letter Number | Ⅼ א | Letterlike numerals like Roman numeral XII | 562 |
| No | Other Number | ½ ① | Fractions, circled digits, superscripts | 915 |
| Pc | Connector Punctuation | _ ‿ | Underscore family, joins words in identifiers | 10 |
| Pd | Dash Punctuation | - – — | Hyphens and dashes of every width | 27 |
| Ps | Open Punctuation | ( [ { | Opening brackets | 80 |
| Pe | Close Punctuation | ) ] } | Closing brackets | 78 |
| Pi | Initial Quote | « “ | Opening quotation marks | 12 |
| Pf | Final Quote | » ” | Closing quotation marks | 10 |
| Po | Other Punctuation | ! ? ; | Everything punctuational not in a P subclass | 643 |
| Sm | Math Symbol | + = ∞ | Operators from plus to infinity | 1,005 |
| Sc | Currency Symbol | $ € ¥ | The money glyphs | 67 |
| Sk | Modifier Symbol | ^ ¸ | Circumflex, cedilla keys - phonetic modifiers | 127 |
| So | Other Symbol | © ✓ ☺ | Symbols that are neither math nor currency nor modifier | 7,561 |
| Zs | Space Separator | space NBSP | Spaces of every width | 17 |
| Zl | Line Separator | U+2028 | The one explicit line separator | 1 |
| Zp | Paragraph Separator | U+2029 | The one explicit paragraph separator | 1 |
| Cc | Control | U+0009 tab U+001B esc | C0/C1 controls: tab, escape, bell | 65 |
| Cf | Format | ZWJ U+00ad | Invisible directives: zero-width joiner, soft hyphen, bidi marks | 170 |
| Cs | Surrogate | U+D800-U+DFFF | UTF-16 plumbing, invalid as characters | 6 |
| Co | Private Use | U+E000+ | Vendor-assigned: emoji pre-2010, game glyphs | 6 |
| Cn | Unassigned | — | Reserved slots awaiting future characters | 1,114,112 minus assigned |
Before Unicode lets any tool work with a character, it answers one question: what IS this? The answer is the general category (GC) - one of 30 codes like Lu (uppercase letter), Nd (decimal digit), Sm (math symbol) or Cn (unassigned). This table lists all 30 with examples, plain-English descriptions and live counts from the Unicode Character Database.
The categories are the load-bearing property of the standard: regex engines translate \p{Lu} into them, word counters count L* vs punctuation, password validators reject Cc controls, and font shapers consult M* marks before placing accents. When two characters look alike but behave differently, the category is usually the reason.
How to use
- Type a code, name or example to filter - 'Lu', 'digit', 'mark' and 'separator' all find their rows.
- Click any code to copy the two-letter abbreviation for use in regex \p{...} patterns.
- Read the counts column to see where Unicode's mass lives - the letter buckets dwarf everything else combined.
Frequently asked questions
Why does Lo have more characters than every other category combined?
Lo (other letters) is where Unicode puts every writing system's ordinary letters that are not case-mapped: the ~21,000 CJK ideographs, all the kana, Hangul syllables, Arabic, Hebrew, Devanagari and hundreds more scripts. Upper and lower case are a European inheritance - Lu and Ll only make sense for bicameral alphabets like Latin, Greek and Cyrillic, so the world's majority scripts all funnel into Lo.
What is the difference between Nd, Nl and No?
Nd (decimal numbers) are the digit characters that can build run-of-the-mill numbers - 0-9 plus the digits of other scripts like Devanagari २. Nl (letter numbers) are letterlike numerals such as the Roman numeral Ⅼ. No (other numbers) are everything else numeric - fractions like ½, circled digits like ①, superscripts like ². The split matters in code: parseInt accepts Nd digits from any script, but a ² in a form field is No, not a digit.
Why do Mn combining marks matter in validators and counters?
Mn (nonspacing marks) are accents that stack onto the previous character - the dot of an i when written decomposed is a separate Mn code point. A naive length check counts e + ́ as two characters; a grapheme-aware one counts one letter. The same trap hits regex character classes: [a-z] silently excludes decomposed accents, while \p{M} is how you match the marks deliberately.
What lives in Cf, the invisible category?
Cf (format) characters change how neighboring text behaves while occupying no visual space: the zero-width joiner that stitches emoji sequences, the soft hyphen that marks legal break points, and the bidirectional marks that reorder Arabic and Hebrew runs. They are also the favorite tool of homograph spoofing - invisible characters hidden inside lookalike domains - which is why security filters strip or flag Cf before display.