Unicode Categories Table

–click a code to copy it
CodeCategoryExamplesWhat belongs hereAssigned
LuUppercase LetterA É ΘCapital letters across all alphabets1,906
LlLowercase Lettera é θSmall letters across all alphabets2,366
LtTitlecase LetterDž LjDigraph letters like Dz that title-case half31
LmModifier Letterˆ ʰLetters that modify the next sound, not standalone words473
LoOther Letter中 あ אEvery ideograph, kana, abjad letter - the biggest bucket21,852
MnNonspacing Marḱ ̈Accents that overlay the previous character2,090
McSpacing Markः ัCombining marks that claim their own width477
MeEnclosing Mark⃝Rings and circles that wrap the previous character13
NdDecimal Number0 3 २Digits usable in decimal numbers, any script770
NlLetter NumberⅬ אLetterlike numerals like Roman numeral XII562
NoOther Number½ ①Fractions, circled digits, superscripts915
PcConnector Punctuation_ ‿Underscore family, joins words in identifiers10
PdDash Punctuation- – —Hyphens and dashes of every width27
PsOpen Punctuation( [ {Opening brackets80
PeClose Punctuation) ] }Closing brackets78
PiInitial Quote« “Opening quotation marks12
PfFinal Quote» ”Closing quotation marks10
PoOther Punctuation! ? ;Everything punctuational not in a P subclass643
SmMath Symbol+ = ∞Operators from plus to infinity1,005
ScCurrency Symbol$ € ¥The money glyphs67
SkModifier Symbol^ ¸Circumflex, cedilla keys - phonetic modifiers127
SoOther Symbol© ✓ ☺Symbols that are neither math nor currency nor modifier7,561
ZsSpace Separatorspace NBSPSpaces of every width17
ZlLine SeparatorU+2028The one explicit line separator1
ZpParagraph SeparatorU+2029The one explicit paragraph separator1
CcControlU+0009 tab U+001B escC0/C1 controls: tab, escape, bell65
CfFormatZWJ U+00adInvisible directives: zero-width joiner, soft hyphen, bidi marks170
CsSurrogateU+D800-U+DFFFUTF-16 plumbing, invalid as characters6
CoPrivate UseU+E000+Vendor-assigned: emoji pre-2010, game glyphs6
CnUnassigned—Reserved slots awaiting future characters1,114,112 minus assigned
The 30 general categories are the first split Unicode performs on every code point - the GC property in the Unicode Character Database - and regex engines, word counters and password validators all branch on them. The counts column is computed live from the current UCD: Lo (other letters) dominates with over 21,000 because CJK ideographs and every abjad land there. Bottom line: is it a letter, a number, a mark or a symbol? That four-way question is what these codes answer, and two characters that look identical can sit in different categories - the superscript ² is No while a plain 2 is Nd, which is exactly why form validators should reject No digits. Inspect companions: Unicode character inspector, Unicode blocks table, math symbols table and script codes table.

Before Unicode lets any tool work with a character, it answers one question: what IS this? The answer is the general category (GC) - one of 30 codes like Lu (uppercase letter), Nd (decimal digit), Sm (math symbol) or Cn (unassigned). This table lists all 30 with examples, plain-English descriptions and live counts from the Unicode Character Database.

The categories are the load-bearing property of the standard: regex engines translate \p{Lu} into them, word counters count L* vs punctuation, password validators reject Cc controls, and font shapers consult M* marks before placing accents. When two characters look alike but behave differently, the category is usually the reason.

How to use

  1. Type a code, name or example to filter - 'Lu', 'digit', 'mark' and 'separator' all find their rows.
  2. Click any code to copy the two-letter abbreviation for use in regex \p{...} patterns.
  3. Read the counts column to see where Unicode's mass lives - the letter buckets dwarf everything else combined.

Frequently asked questions

Why does Lo have more characters than every other category combined?

Lo (other letters) is where Unicode puts every writing system's ordinary letters that are not case-mapped: the ~21,000 CJK ideographs, all the kana, Hangul syllables, Arabic, Hebrew, Devanagari and hundreds more scripts. Upper and lower case are a European inheritance - Lu and Ll only make sense for bicameral alphabets like Latin, Greek and Cyrillic, so the world's majority scripts all funnel into Lo.

What is the difference between Nd, Nl and No?

Nd (decimal numbers) are the digit characters that can build run-of-the-mill numbers - 0-9 plus the digits of other scripts like Devanagari २. Nl (letter numbers) are letterlike numerals such as the Roman numeral Ⅼ. No (other numbers) are everything else numeric - fractions like ½, circled digits like ①, superscripts like ². The split matters in code: parseInt accepts Nd digits from any script, but a ² in a form field is No, not a digit.

Why do Mn combining marks matter in validators and counters?

Mn (nonspacing marks) are accents that stack onto the previous character - the dot of an i when written decomposed is a separate Mn code point. A naive length check counts e + ́ as two characters; a grapheme-aware one counts one letter. The same trap hits regex character classes: [a-z] silently excludes decomposed accents, while \p{M} is how you match the marks deliberately.

What lives in Cf, the invisible category?

Cf (format) characters change how neighboring text behaves while occupying no visual space: the zero-width joiner that stitches emoji sequences, the soft hyphen that marks legal break points, and the bidirectional marks that reorder Arabic and Hebrew runs. They are also the favorite tool of homograph spoofing - invisible characters hidden inside lookalike domains - which is why security filters strip or flag Cf before display.

Related tools