UTF-8 is a variable-width encoding that represents any Unicode character as a sequence of 1 to 4 bytes. It was defined in RFC 3629 by the Internet Engineering Task Force. The first 128 code points (standard ASCII) use exactly one byte each, which is why UTF-8 is fully backward-compatible with ASCII.
How the Byte Patterns Work
UTF-8 uses a clever scheme of leading bits to signal how many bytes a character occupies:
| Code point range | Bytes | Byte 1 | Byte 2 | Byte 3 | Byte 4 |
|---|---|---|---|---|---|
| U+0000 – U+007F | 1 | 0xxxxxxx | |||
| U+0080 – U+07FF | 2 | 110xxxxx | 10xxxxxx | ||
| U+0800 – U+FFFF | 3 | 1110xxxx | 10xxxxxx | 10xxxxxx | |
| U+10000 – U+10FFFF | 4 | 11110xxx | 10xxxxxx | 10xxxxxx | 10xxxxxx |
The x positions carry the actual bits of the code point. Leading bytes start with 0, 110, 1110, or 11110. Continuation bytes always start with 10, so software can tell immediately whether a byte is a character start or a continuation.
Worked Example: The Letter é
The accented letter é has Unicode code point U+00E9 (decimal 233).
- 233 falls in the range U+0080–U+07FF, so é uses 2 bytes.
- The code point in binary is
11101001. - Split into the UTF-8 template
110xxxxx 10xxxxxx:110+00011=11000011= hexC310+101001=10101001= hexA9
- UTF-8 bytes: C3 A9
You can verify this: Buffer.from('é', 'utf8').toString('hex') in Node.js returns c3a9.
Four Real Examples
| Character | Name | Code point | UTF-8 bytes | Byte count |
|---|---|---|---|---|
| A | Latin A | U+0041 | 41 | 1 |
| é | Latin e acute | U+00E9 | C3 A9 | 2 |
| € | Euro sign | U+20AC | E2 82 AC | 3 |
| 😀 | Grinning face | U+1F600 | F0 9F 98 80 | 4 |
The euro sign (€) at U+20AC requires 3 bytes because it falls in the U+0800–U+FFFF range. The grinning face emoji at U+1F600 requires 4 bytes because it is above U+FFFF.
Why UTF-8 Won
Before UTF-8, software used dozens of incompatible encodings (Latin-1, Shift-JIS, Windows-1252, and others). A file encoded in Latin-1 would display garbled text when opened as Shift-JIS. UTF-8 ended that fragmentation for the web.
Key advantages:
- ASCII-safe: files that contain only ASCII bytes are valid UTF-8. No conversion needed.
- Self-synchronizing: you can start reading anywhere in a stream and immediately tell whether you are at a character boundary.
- Compact: English text stays at one byte per character. Emoji and rare scripts pay more, but common text costs nothing extra.
RFC 3629, which standardized UTF-8 in 2003, explicitly deprecated all other UTF-8 variants and established it as the preferred encoding for internet protocols.
Practical Takeaways
- Always declare
charset=UTF-8in HTML documents. - Save source code files in UTF-8 to avoid “mojibake” (garbled text from encoding mismatch).
- String length in bytes and string length in characters are not the same.
strlen('é')is 2 in a byte-counting function but 1 in a character-counting one — a common source of bugs in web form validation.
UTF-8 vs. UTF-16 vs. UTF-32
All three encode the same Unicode code points, but differently:
| Encoding | Byte width | Common use |
|---|---|---|
| UTF-8 | 1–4 bytes/char | Web, Linux, macOS, most text files |
| UTF-16 | 2 or 4 bytes/char | Windows internally, Java, JavaScript strings |
| UTF-32 | Always 4 bytes | Simple processing where fixed width helps |
UTF-16 uses 2 bytes for most characters and 4 bytes (a surrogate pair) for code points above U+FFFF. Java’s char type is UTF-16, which means a single emoji requires two char values — a common source of bugs in Java string length calculations.
UTF-32 is wasteful for English (4 bytes per character, vs. 1 for ASCII in UTF-8) but simple to process because every character takes the same number of bytes, making character indexing trivial.
UTF-8 wins on the web because it is compact, self-synchronizing, and ASCII-compatible. RFC 3629 established it as the required encoding for new IETF protocols in 2003.
See ASCII vs. Unicode to understand the full character standard, or use the ASCII converter to inspect the code point and byte representation of any character.