Skip to content
CodeShift
Menu

What Is UTF-8? Encoding Explained

UTF-8 encodes any Unicode character as 1 to 4 bytes. ASCII uses 1 byte; é uses 2 bytes (C3 A9); emoji use 4 bytes. It is the encoding of the modern web.

By The CodeShift DeskPublished September 8, 2026

UTF-8 is a variable-width encoding that represents any Unicode character as a sequence of 1 to 4 bytes. It was defined in RFC 3629 by the Internet Engineering Task Force. The first 128 code points (standard ASCII) use exactly one byte each, which is why UTF-8 is fully backward-compatible with ASCII.

How the Byte Patterns Work

UTF-8 uses a clever scheme of leading bits to signal how many bytes a character occupies:

Code point range Bytes Byte 1 Byte 2 Byte 3 Byte 4
U+0000 – U+007F 1 0xxxxxxx
U+0080 – U+07FF 2 110xxxxx 10xxxxxx
U+0800 – U+FFFF 3 1110xxxx 10xxxxxx 10xxxxxx
U+10000 – U+10FFFF 4 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx

The x positions carry the actual bits of the code point. Leading bytes start with 0, 110, 1110, or 11110. Continuation bytes always start with 10, so software can tell immediately whether a byte is a character start or a continuation.

Worked Example: The Letter é

The accented letter é has Unicode code point U+00E9 (decimal 233).

  • 233 falls in the range U+0080–U+07FF, so é uses 2 bytes.
  • The code point in binary is 11101001.
  • Split into the UTF-8 template 110xxxxx 10xxxxxx:
    • 110 + 00011 = 11000011 = hex C3
    • 10 + 101001 = 10101001 = hex A9
  • UTF-8 bytes: C3 A9

You can verify this: Buffer.from('é', 'utf8').toString('hex') in Node.js returns c3a9.

Four Real Examples

Character Name Code point UTF-8 bytes Byte count
A Latin A U+0041 41 1
é Latin e acute U+00E9 C3 A9 2
€ Euro sign U+20AC E2 82 AC 3
😀 Grinning face U+1F600 F0 9F 98 80 4

The euro sign (€) at U+20AC requires 3 bytes because it falls in the U+0800–U+FFFF range. The grinning face emoji at U+1F600 requires 4 bytes because it is above U+FFFF.

Why UTF-8 Won

Before UTF-8, software used dozens of incompatible encodings (Latin-1, Shift-JIS, Windows-1252, and others). A file encoded in Latin-1 would display garbled text when opened as Shift-JIS. UTF-8 ended that fragmentation for the web.

Key advantages:

  • ASCII-safe: files that contain only ASCII bytes are valid UTF-8. No conversion needed.
  • Self-synchronizing: you can start reading anywhere in a stream and immediately tell whether you are at a character boundary.
  • Compact: English text stays at one byte per character. Emoji and rare scripts pay more, but common text costs nothing extra.

RFC 3629, which standardized UTF-8 in 2003, explicitly deprecated all other UTF-8 variants and established it as the preferred encoding for internet protocols.

Practical Takeaways

  • Always declare charset=UTF-8 in HTML documents.
  • Save source code files in UTF-8 to avoid “mojibake” (garbled text from encoding mismatch).
  • String length in bytes and string length in characters are not the same. strlen('é') is 2 in a byte-counting function but 1 in a character-counting one — a common source of bugs in web form validation.

UTF-8 vs. UTF-16 vs. UTF-32

All three encode the same Unicode code points, but differently:

Encoding Byte width Common use
UTF-8 1–4 bytes/char Web, Linux, macOS, most text files
UTF-16 2 or 4 bytes/char Windows internally, Java, JavaScript strings
UTF-32 Always 4 bytes Simple processing where fixed width helps

UTF-16 uses 2 bytes for most characters and 4 bytes (a surrogate pair) for code points above U+FFFF. Java’s char type is UTF-16, which means a single emoji requires two char values — a common source of bugs in Java string length calculations.

UTF-32 is wasteful for English (4 bytes per character, vs. 1 for ASCII in UTF-8) but simple to process because every character takes the same number of bytes, making character indexing trivial.

UTF-8 wins on the web because it is compact, self-synchronizing, and ASCII-compatible. RFC 3629 established it as the required encoding for new IETF protocols in 2003.

See ASCII vs. Unicode to understand the full character standard, or use the ASCII converter to inspect the code point and byte representation of any character.

Frequently asked questions

How many bytes does UTF-8 use per character?+

Between 1 and 4 bytes depending on the code point. ASCII characters use 1 byte. Most Western European accented characters use 2. Chinese, Japanese, and Korean characters use 3. Emoji typically use 4.

Is UTF-8 the same as Unicode?+

No. Unicode is the standard that assigns code points to characters. UTF-8 is one way to encode those code points as bytes. UTF-16 and UTF-32 are other encodings of the same Unicode code points.

What happens if I open a UTF-8 file as ASCII?+

Characters within the 0–127 range display correctly. Any character that uses more than one byte will appear as garbage characters or boxes, because the extra bytes are not valid ASCII.

Do I need a BOM (byte order mark) with UTF-8?+

No. RFC 3629 advises against using a BOM in UTF-8. Unlike UTF-16, UTF-8 has no byte-order ambiguity. A BOM at the start of a UTF-8 file can cause problems in some programs.

What encoding does HTML use by default?+

HTML5 defaults to UTF-8. You declare it with the meta tag: meta charset equals UTF-8. Without this, browsers may guess the encoding incorrectly.

Keep reading