ASCII is a 128-character set that covers the English alphabet, digits, and basic punctuation — everything you can type on a standard US keyboard. Unicode is a universal standard that assigns a unique number to every character in every human writing system. The two are compatible: the first 128 Unicode characters are exactly the ASCII set.
ASCII: The Basics
The American Standard Code for Information Interchange was first published in 1963 and finalized as ANSI X3.4 in 1968. It maps 128 characters to numbers 0–127 using 7 bits per character.
| Range | Characters | Examples |
|---|---|---|
| 0–31 | Control codes (non-printable) | Line feed, tab |
| 32–47 | Punctuation and space | Space, !, “ |
| 48–57 | Digits | 0–9 |
| 65–90 | Uppercase letters | A–Z |
| 97–122 | Lowercase letters | a–z |
The full table is at /ascii-table/. One byte stores one ASCII character; the eighth bit was unused and later reused for “extended ASCII” in regional variants — a source of endless incompatibility.
Unicode: One Standard for All Scripts
The Unicode Consortium released Unicode 1.0 in 1991 to solve that incompatibility. Unicode assigns each character a code point — a unique number written as U+ followed by a hex value. The uppercase letter A is U+0041; the Chinese character 中 is U+4E2D; the grinning face emoji is U+1F600.
Unicode currently defines over 149,000 characters across scripts including Latin, Arabic, Devanagari, Chinese, Japanese, Korean, Hebrew, and dozens of historical scripts.
Comparison Table
| Feature | ASCII | Unicode |
|---|---|---|
| First published | 1963 | 1991 |
| Character count | 128 | 149,813+ (as of Unicode 15.1) |
| Bit width | 7 bits per character | Variable (code points up to 21 bits) |
| Scripts covered | Latin (English only) | 161+ scripts |
| Encoding formats | Fixed 7-bit | UTF-8, UTF-16, UTF-32, and others |
| Web standard? | No (superseded) | Yes (UTF-8) |
How They Relate
ASCII is a subset of Unicode. Code points U+0000–U+007F correspond exactly to the 128 ASCII characters. Text written in pure ASCII is valid Unicode with no conversion needed.
The difference appears beyond those 128 code points. The accented letter é is U+00E9 — code point 233, outside ASCII’s range. To store it, you need a Unicode encoding such as UTF-8, which represents é as two bytes: C3 A9.
Which Encoding to Choose
For any new project: use UTF-8. It is the default encoding for HTML5, JSON, and most modern programming languages. It is backward-compatible with ASCII, compact for English text (each ASCII character uses one byte), and can represent any Unicode character. Web crawl surveys consistently show UTF-8 used on the overwhelming majority of publicly accessible web pages.
Extended ASCII and the Compatibility Problem
The original ASCII left the eighth bit of each byte unused. Hardware vendors filled it with 128 additional characters for regional needs: Western European languages used ISO 8859-1 (Latin-1), adding é, ü, ñ, and similar characters. Central European languages used ISO 8859-2. Cyrillic text used ISO 8859-5. The result was dozens of incompatible “extended ASCII” variants — the same byte value meant different characters in different countries.
This fragmentation caused what Japanese programmers called “mojibake” — garbled characters produced when software assumed the wrong encoding. An email written in Latin-1 read as Cyrillic would display nonsense. Unicode resolved this by defining a single global table where each character has one number regardless of country or language.
Unicode Code Points in Practice
Every character in Unicode has a code point, written as U+ followed by 4 to 6 hex digits. Common examples:
| Character | Name | Code point | Category |
|---|---|---|---|
| A | Latin capital A | U+0041 | Letter |
| é | Latin e with acute | U+00E9 | Letter |
| © | Copyright sign | U+00A9 | Symbol |
| € | Euro sign | U+20AC | Currency |
| ♥ | Black heart suit | U+2665 | Symbol |
| 中 | CJK character | U+4E2D | Letter |
| 😀 | Grinning face | U+1F600 | Emoji |
Code points are not bytes — they are abstract numbers. How those numbers become bytes in a file depends on the encoding (UTF-8, UTF-16, UTF-32). See What Is UTF-8? for how the encoding works byte by byte, or use the ASCII converter to translate characters to their numeric codes and back.