EBCDIC Table

–CP037 rows verified
Character runEBCDIC (CP037)ASCIIWhy it sits there
Space4020the most-reliable fingerprint: EBCDIC space is 0x40, ASCII space is 0x20
Digits 0-9F0-F930-39digit nibble is F - the BCD zone inheritance
Lowercase a-i81-8961-69first of three lowercase runs
Lowercase j-r91-996A-72second run - a 7-slot gap sits between runs
Lowercase s-zA2-A973-7Athird run - s deliberately opens on 2, not 1
Uppercase A-IC1-C941-49first uppercase run, zone nibble C
Uppercase J-RD1-D94A-52second run, zone nibble D
Uppercase S-ZE2-E953-5Athird run, zone nibble E
CharASCIICP037CP037 decCharacter name
!0x210x5A90Exclamation Mark
"0x220x7F127Quotation Mark
#0x230x7B123Number Sign
$0x240x5B91Dollar Sign
%0x250x6C108Percent Sign
&0x260x5080Ampersand
'0x270x7D125Apostrophe
(0x280x4D77Left Parenthesis
)0x290x5D93Right Parenthesis
*0x2A0x5C92Asterisk
+0x2B0x4E78Plus Sign
,0x2C0x6B107Comma
-0x2D0x6096Hyphen-Minus
.0x2E0x4B75Full Stop
/0x2F0x6197Solidus
:0x3A0x7A122Colon
;0x3B0x5E94Semicolon
<0x3C0x4C76Less-Than Sign
=0x3D0x7E126Equals Sign
>0x3E0x6E110Greater-Than Sign
?0x3F0x6F111Question Mark
@0x400x7C124Commercial At
[0x5B0xBA186Left Square Bracket
\0x5C0xE0224Reverse Solidus
]0x5D0xBB187Right Square Bracket
^0x5E0xB0176Circumflex Accent
_0x5F0x6D109Low Line
`0x600x79121Grave Accent
{0x7B0xC0192Left Curly Bracket
|0x7C0x4F79Vertical Line
}0x7D0xD0208Right Curly Bracket
~0x7E0xA1161Tilde
The structure chart is the history: EBCDIC kept the punched-card column layout its ancestor coded - each case of letters split into three runs (zones 8/9/A lowercase, C/D/E uppercase) with gaps between, and digits keeping their BCD zone nibble F - which is why the arithmetic ASCII programmers take for granted breaks: Wikipedia's compatibility section notes that the C loop for (c = 'A'; c <= 'Z'; ++c) putchar(c); “would print the letters from A to Z if ASCII is used, but print 41 characters” on EBCDIC. The punctuation chart shows the other famous cost: the Jargon File says EBCDIC “exists in at least six mutually incompatible versions, all featuring such delights as non-contiguous letter sequences and the absence of several ASCII punctuation characters fairly important for modern computer languages” - in CP037 the brackets [] live at 0xBA/0xBB and the braces {} at 0xC0/0xD0.
Authority chain: the values above are generated from the official Microsoft/Unicode mapping (CP037.TXT, “cp037_IBMUSCanada to Unicode table”) and byte-verified, not transcribed. EBCDIC was “devised in 1963 and 1964 by IBM and announced with the release of the IBM System/360”, descended from punched-card six-bit BCD code - and it still runs z/OS mainframes in banking, airlines and government, while IBM AIX, Linux on IBM Z and Linux on Power all use ASCII. One newline trap: CP037 byte 0x15 is NL (Unicode NEXT LINE, U+0085), not the ASCII line feed. Byte-side kin on this site: the ASCII table for the 128-character reference, text-to-binary for either encoding bit for bit, and the Windows-1251 table for the Cyrillic single-byte world.

EBCDIC - Extended Binary Coded Decimal Interchange Code - is the eight-bit character encoding Wikipedia says is used mainly on IBM mainframe and IBM midrange computer operating systems, devised in 1963 and 1964 by IBM and announced with the release of the IBM System/360 line of mainframe computers. It is eight bits developed separately from seven-bit ASCII, and its family tree explains every weird number in the chart below: it descended from the code used with punched cards and the corresponding six-bit binary-coded decimal code used with most of IBM’s computer peripherals of the late 1950s and early 1960s.

The first chart shows the consequence: letters run in three separate segments per case (a-i, j-r, s-z with zone nibbles 8, 9, A) and digits keep their BCD zone (0-9 lives at 0xF0-0xF9, upper nibble F), while ASCII puts each alphabet in one contiguous run. The second chart maps all 31 ASCII punctuation characters into CP037, the US/Canada EBCDIC code page - and shows how far punctuation travels: the exclamation mark lands at 0x5A, the dollar sign at 0x5B, the quotation mark all the way up at 0x7F. Space is 0x40 in EBCDIC versus 0x20 in ASCII - the single most reliable fingerprint when identifying unknown files.

Every value here is generated from Microsoft and Unicode’s published CP037 mapping and byte-verified, not transcribed. The encoding matters in 2026 because z/OS mainframes still run banking, airline and government batch workloads natively in EBCDIC - while IBM’s own AIX and Linux on Z use ASCII, so the boundary is real and file-level. For the ASCII side of every row, the ASCII table on this site lists the full 128-character set, and the text-to-binary converter shows either encoding bit for bit.

How to use

  1. Identify an unknown file in two bytes: open it in a hex viewer and look at where spaces (0x40 vs 0x20) and capital letters (0xC1-0xE9 vs 0x41-0x5A) fall. Both charts above give you the byte-level fingerprint; if letters cluster in the C-E zone, you are looking at EBCDIC.
  2. Convert field data using the punctuation chart: legacy mainframe reports and COBOL-era files use CP037 punctuation positions (for example, decimal points at 0x4B and commas at 0x6B). Mapping through this table restores the ASCII positions before any text processing.
  3. Never assume one EBCDIC: confirm the exact code page (037 for US/Canada, 500 international, 1047 on z/OS Unix services) before bulk conversion - the variants disagree on exactly the punctuation and bracket characters scripts depend on. The official CP037 mapping this chart is generated from is published by Microsoft and Unicode.

Frequently asked questions

Why are EBCDIC letters not contiguous?

Because the letters were laid down as three separate runs per case, each in a different zone nibble: a-i in 0x81-0x89, j-r in 0x91-0x99, s-z in 0xA2-0xAA (uppercase the same with zones C, D, E). The chart above lists every run with its ASCII counterpart. The practical casualty is arithmetic: Wikipedia’s compatibility section notes that the C loop for (c = 'A'; c <= 'Z'; ++c) putchar(c); would print the letters from A to Z if ASCII is used, but print 41 characters on EBCDIC - the gaps between runs get traversed too.

Is EBCDIC still used in 2026?

Yes, in a specific fortress: IBM z/OS mainframes - the systems running core banking, airline reservations and government batch workloads - remain EBCDIC-native. Wikipedia is precise about the boundary though: IBM AIX, Linux on IBM Z and Linux on Power all use ASCII, as does everything running on the IBM PC line. So EBCDIC lives inside classic z/OS subsystems and the file interfaces around them, and almost nowhere else.

How do I tell whether a file is EBCDIC or ASCII?

Check the spaces and the letters. A text file whose every space is byte 0x40 is EBCDIC (ASCII space is 0x20); a file whose letters concentrate in 0xC1-0xE9 is EBCDIC uppercase (ASCII uppercase lives at 0x41-0x5A). Punctuation is the third tell - in CP037 the exclamation mark is 0x5A and the dollar sign 0x5B, positions that hold entirely different characters in ASCII, as the punctuation chart above shows byte for byte.

How many EBCDIC variants are there?

More than anyone wants: the Jargon File’s entry says it exists in at least six mutually incompatible versions, all featuring such delights as non-contiguous letter sequences and the absence of several ASCII punctuation characters fairly important for modern computer languages. The famous ones are code pages 037 (US/Canada - the one charted here), 500 (international), 1047 (Unix services on z/OS) and the euro-era 1140-family. Always confirm the exact code page before converting.

What is the newline situation in EBCDIC?

Different from ASCII by design: CP037 maps byte 0x15 to NL, which Unicode records as NEXT LINE (NEL, U+0085) - not the ASCII line feed U+000A. Microsoft’s CP037 mapping and Unicode’s documentation both flag the mismatch: translating EBCDIC to ASCII code-for-code will not produce Unix-style line breaks, which is why mainframe file transfers have their own newline conventions.

Related tools