Skip to content
CodeShift
Menu

ASCII vs. Unicode: What's the Difference?

ASCII covers 128 characters for English. Unicode covers over 149,000 characters for every human writing system. UTF-8 encodes both and is the web standard.

By The CodeShift DeskPublished September 5, 2026

ASCII is a 128-character set that covers the English alphabet, digits, and basic punctuation — everything you can type on a standard US keyboard. Unicode is a universal standard that assigns a unique number to every character in every human writing system. The two are compatible: the first 128 Unicode characters are exactly the ASCII set.

ASCII: The Basics

The American Standard Code for Information Interchange was first published in 1963 and finalized as ANSI X3.4 in 1968. It maps 128 characters to numbers 0–127 using 7 bits per character.

Range Characters Examples
0–31 Control codes (non-printable) Line feed, tab
32–47 Punctuation and space Space, !, “
48–57 Digits 0–9
65–90 Uppercase letters A–Z
97–122 Lowercase letters a–z

The full table is at /ascii-table/. One byte stores one ASCII character; the eighth bit was unused and later reused for “extended ASCII” in regional variants — a source of endless incompatibility.

Unicode: One Standard for All Scripts

The Unicode Consortium released Unicode 1.0 in 1991 to solve that incompatibility. Unicode assigns each character a code point — a unique number written as U+ followed by a hex value. The uppercase letter A is U+0041; the Chinese character 中 is U+4E2D; the grinning face emoji is U+1F600.

Unicode currently defines over 149,000 characters across scripts including Latin, Arabic, Devanagari, Chinese, Japanese, Korean, Hebrew, and dozens of historical scripts.

Comparison Table

Feature ASCII Unicode
First published 1963 1991
Character count 128 149,813+ (as of Unicode 15.1)
Bit width 7 bits per character Variable (code points up to 21 bits)
Scripts covered Latin (English only) 161+ scripts
Encoding formats Fixed 7-bit UTF-8, UTF-16, UTF-32, and others
Web standard? No (superseded) Yes (UTF-8)

How They Relate

ASCII is a subset of Unicode. Code points U+0000–U+007F correspond exactly to the 128 ASCII characters. Text written in pure ASCII is valid Unicode with no conversion needed.

The difference appears beyond those 128 code points. The accented letter é is U+00E9 — code point 233, outside ASCII’s range. To store it, you need a Unicode encoding such as UTF-8, which represents é as two bytes: C3 A9.

Which Encoding to Choose

For any new project: use UTF-8. It is the default encoding for HTML5, JSON, and most modern programming languages. It is backward-compatible with ASCII, compact for English text (each ASCII character uses one byte), and can represent any Unicode character. Web crawl surveys consistently show UTF-8 used on the overwhelming majority of publicly accessible web pages.

Extended ASCII and the Compatibility Problem

The original ASCII left the eighth bit of each byte unused. Hardware vendors filled it with 128 additional characters for regional needs: Western European languages used ISO 8859-1 (Latin-1), adding é, ü, ñ, and similar characters. Central European languages used ISO 8859-2. Cyrillic text used ISO 8859-5. The result was dozens of incompatible “extended ASCII” variants — the same byte value meant different characters in different countries.

This fragmentation caused what Japanese programmers called “mojibake” — garbled characters produced when software assumed the wrong encoding. An email written in Latin-1 read as Cyrillic would display nonsense. Unicode resolved this by defining a single global table where each character has one number regardless of country or language.

Unicode Code Points in Practice

Every character in Unicode has a code point, written as U+ followed by 4 to 6 hex digits. Common examples:

Character Name Code point Category
A Latin capital A U+0041 Letter
é Latin e with acute U+00E9 Letter
© Copyright sign U+00A9 Symbol
€ Euro sign U+20AC Currency
♥ Black heart suit U+2665 Symbol
中 CJK character U+4E2D Letter
😀 Grinning face U+1F600 Emoji

Code points are not bytes — they are abstract numbers. How those numbers become bytes in a file depends on the encoding (UTF-8, UTF-16, UTF-32). See What Is UTF-8? for how the encoding works byte by byte, or use the ASCII converter to translate characters to their numeric codes and back.

Frequently asked questions

Is ASCII part of Unicode?+

Yes. The first 128 Unicode code points (U+0000 through U+007F) are identical to ASCII. Any ASCII text is valid UTF-8 without modification.

Why did we need Unicode if ASCII worked?+

ASCII only handles English. Accented characters, non-Latin scripts (Chinese, Arabic, Hindi), and symbols like the euro sign all require more than 128 code points.

What is the difference between Unicode and UTF-8?+

Unicode is the standard that assigns a code point number to every character. UTF-8 is one encoding that stores those numbers in memory as bytes. Others include UTF-16 and UTF-32.

How many characters does Unicode support?+

Unicode 15.1 defines 149,813 characters across 161 scripts, including emoji. The standard can accommodate over 1.1 million code points in total.

What encoding should I use for a new application?+

Use UTF-8. It is the dominant encoding on the web, compatible with ASCII, and handles every Unicode character. The Internet Engineering Task Force requires it as the default for new internet protocols.

Keep reading