Windows-1252 Table

–bytes in the divergent range
ByteCharUnicodeDecUTF-8Latin-1 counterpart
80€U+20AC Euro Sign128E2 82 ACC1 control (Padding Character) - the Latin-1 slot Windows replaced
81–undefined129–unassigned (5 such slots); browsers substitute the Windows glyph or nothing
82‚U+201A Single Low-9 Quotation Mark130E2 80 9AC1 control (Break Permitted Here) - the Latin-1 slot Windows replaced
83ƒU+0192 Latin Small Letter F With Hook131C6 92C1 control (No Break Here) - the Latin-1 slot Windows replaced
84„U+201E Double Low-9 Quotation Mark132E2 80 9EC1 control (Index) - the Latin-1 slot Windows replaced
85…U+2026 Horizontal Ellipsis133E2 80 A6C1 control (Next Line) - the Latin-1 slot Windows replaced
86†U+2020 Dagger134E2 80 A0C1 control (Start Of Selected Area) - the Latin-1 slot Windows replaced
87‡U+2021 Double Dagger135E2 80 A1C1 control (End Of Selected Area) - the Latin-1 slot Windows replaced
88ˆU+02C6 Modifier Letter Circumflex Accent136CB 86C1 control (Character Tabulation Set) - the Latin-1 slot Windows replaced
89‰U+2030 Per Mille Sign137E2 80 B0C1 control (Character Tabulation With Justification) - the Latin-1 slot Windows replaced
8AŠU+0160 Latin Capital Letter S With Caron138C5 A0C1 control (Line Tabulation Set) - the Latin-1 slot Windows replaced
8B‹U+2039 Single Left-Pointing Angle Quotation Mark139E2 80 B9C1 control (Partial Line Forward) - the Latin-1 slot Windows replaced
8CŒU+0152 Latin Capital Ligature Oe140C5 92C1 control (Partial Line Backward) - the Latin-1 slot Windows replaced
8D–undefined141–unassigned (5 such slots); browsers substitute the Windows glyph or nothing
8EŽU+017D Latin Capital Letter Z With Caron142C5 BDC1 control (Single Shift Two) - the Latin-1 slot Windows replaced
8F–undefined143–unassigned (5 such slots); browsers substitute the Windows glyph or nothing
90–undefined144–unassigned (5 such slots); browsers substitute the Windows glyph or nothing
91‘U+2018 Left Single Quotation Mark145E2 80 98C1 control (Private Use One) - the Latin-1 slot Windows replaced
92’U+2019 Right Single Quotation Mark146E2 80 99C1 control (Private Use Two) - the Latin-1 slot Windows replaced
93“U+201C Left Double Quotation Mark147E2 80 9CC1 control (Set Transmit State) - the Latin-1 slot Windows replaced
94”U+201D Right Double Quotation Mark148E2 80 9DC1 control (Cancel Character) - the Latin-1 slot Windows replaced
95•U+2022 Bullet149E2 80 A2C1 control (Message Waiting) - the Latin-1 slot Windows replaced
96–U+2013 En Dash150E2 80 93C1 control (Start Of Guarded Area) - the Latin-1 slot Windows replaced
97—U+2014 Em Dash151E2 80 94C1 control (End Of Guarded Area) - the Latin-1 slot Windows replaced
98˜U+02DC Small Tilde152CB 9CC1 control (Start Of String) - the Latin-1 slot Windows replaced
99™U+2122 Trade Mark Sign153E2 84 A2C1 control - the Latin-1 slot Windows replaced
9AšU+0161 Latin Small Letter S With Caron154C5 A1C1 control - the Latin-1 slot Windows replaced
9B›U+203A Single Right-Pointing Angle Quotation Mark155E2 80 BAC1 control - the Latin-1 slot Windows replaced
9CœU+0153 Latin Small Ligature Oe156C5 93C1 control - the Latin-1 slot Windows replaced
9D–undefined157–unassigned (5 such slots); browsers substitute the Windows glyph or nothing
9EžU+017E Latin Small Letter Z With Caron158C5 BEC1 control (Single Graphic Character Introducer) - the Latin-1 slot Windows replaced
9FŸU+0178 Latin Capital Letter Y With Diaeresis159C5 B8C1 control (Single Graphic Character Introducer) - the Latin-1 slot Windows replaced
Windows-1252 is the legacy single-byte encoding that Wikipedia calls “the most-used single-byte character encoding in the world” - and, as of September 2026, still the effective encoding of 9.3% of static web pages. Its story is one divergence: “initially the same as ISO 8859-1, it began to diverge starting in Windows 2.0 by adding additional characters in the 0x80 to 0x9F (hex) range” - the 32 slots ISO reserved for C1 control codes, which Microsoft spent on curly quotation marks, dashes, the euro sign and the Å“ of French. This chart lists all 32: 27 printable characters, 5 permanently undefined slots, and the C1 control each byte displaced in real Latin-1. The UTF-8 column is the mojibake key: read ’s UTF-8 bytes (E2 80 99) as Windows-1252 and you get ’ - the classic three-glyph garbage that decodes broken text in one glance.
Two legal facts make this table matter more than its age suggests. First, per the WHATWG Encoding Standard (which Wikipedia cites as the encoding’s governing standard), every modern browser treats a declared ISO 8859-1 as Windows-1252 - the label is a synonym, required by HTML5. Second, the mapping itself is published by Unicode (bestfit1252.txt) and reproduced byte-for-byte by every Python/ICU/WHATWG implementation - the values above were generated from that mapping. Known gaps Wikipedia tabulates: no uppercase ị (not official until 2017), no Slovene č (substituted by č̇), no Dutch Ĭ/ĭ. The escape path is on this site: the UTF-8 encoding table shows where those UTF-8 bytes come from, the smart quotes converter handles the curly-mark problem at the source, and the JS encoding table covers TextDecoder’s windows-1252 label.

Windows-1252 - code page 1252, CP-1252, the encoding Microsoft shipped as the default “ANSI code page” across the Americas, Western Europe, Oceania and much of Africa - is the byte format an entire generation of documents never asked about. Wikipedia’s headline fact: it is “the most-used single-byte character encoding in the world,” and as of September 2026 it is still effectively served on 9.3% of static web pages.

Everything interesting about it lives in 32 bytes. The encoding started as ISO 8859-1 (Latin-1), but from Windows 2.0 Microsoft filled the 0x80–0x9F range - the 32 slots ISO had reserved for unprintable C1 control codes - with the characters Western office work actually needed: curly quotation marks, en and em dashes, the ellipsis, trademark and euro signs, and French œ. Those 27 characters are the whole difference between the two encodings, and this chart lists every one of them.

The table also carries the modern answer key: each byte’s UTF-8 encoding. That column is how you read mojibake - when a UTF-8 apostrophe (E2 80 99) is wrongly decoded as Windows-1252 it becomes ’, and every broken string you have ever pasted from an old export follows the same arithmetic.

How to use

  1. Identify the encoding by its garbage: text that shows ’ where an apostrophe should be, or  before accented letters, is UTF-8 bytes being read as Windows-1252. Find the first garbage glyph in this chart’s char column and the UTF-8 column tells you which bytes produced it.
  2. Use the Dec column for Alt codes: Windows Alt+0146 still produces ’ because 146 decimal is byte 0x92 in this table - the decimal column is the numeric identity of each byte in the divergent range.
  3. Do not declare ISO 8859-1 in new work: the WHATWG Encoding Standard requires browsers to treat that label as Windows-1252, so the declaration is both wrong and ignored. Prefer UTF-8 everywhere - the <a href="https://tooldune.com/utf8-encoding-table/" rel="noopener">UTF-8 encoding table</a> on this site shows its byte grammar.

Frequently asked questions

What is the difference between Windows-1252 and ISO 8859-1?

Exactly 32 bytes. Windows-1252 equals Latin-1 for 0x00&#8211;0x7F and 0xA0&#8211;0xFF, but re-purposes the 32 control-code slots of 0x80&#8211;0x9F for printable characters - curly quotes, dashes, ellipsis, euro - of which 27 are defined and 5 remain unassigned. Wikipedia dates the split to Windows 2.0.

Why do my smart quotes show up as garbage?

Because the document is UTF-8 but something reads it as Windows-1252: a UTF-8 right single quote is the three bytes E2 80 99, and Windows-1252 renders those as &#226;&#8364;&#8482;. Every curly quote, dash and ellipsis in the divergent range produces this three-glyph signature - this table&#8217;s UTF-8 column decodes it.

Is Windows-1252 still used in 2026?

Yes, enough to matter: Wikipedia&#8217;s September 2026 numbers are 9.3% of static web pages (effectively) against only 0.2% declared as Windows-1252 and 1.0% including declared ISO 8859-1 - the gap exists because most sites are served programmatically, and because HTML5 forces browsers to treat a Latin-1 declaration as Windows-1252 anyway. Legacy exports, CSVs and old CMS databases keep it alive.

Does Windows-1252 cover my language?

Western European languages broadly, with famous holes Wikipedia tabulates: no uppercase &#7883; for German (official only since 2017), no Slovene &#269; (substituted with &#269;&#775;), no Dutch &#300;/&#301; pair, and some languages lack their standard quotation marks (German &#8222;quotes&#8220;). If a language needs more than these, it needs Unicode.

What encoding should I use instead?

UTF-8 - &#8220;almost all websites now use the multi-byte character encoding UTF-8, another superset of ASCII,&#8221; per Wikipedia, and every byte in this chart has a UTF-8 equivalent (the table&#8217;s UTF-8 column lists them). New files, APIs and databases should declare UTF-8; the legacy table is for reading old data, not writing new.

Related tools