What Base64 does
Base64 represents any sequence of bytes using only 64 safe, printable characters: A–Z, a–z, 0–9, + and /, with = for padding. It exists so that binary data — images, keys, attachments — can travel through systems built for text, such as email bodies, JSON, data URLs and HTTP headers. The standard reference is RFC 4648.
How the encoding works: “Man” → TWFu
Base64 takes the input three bytes (24 bits) at a time and splits them into four 6-bit groups. Each 6-bit value, 0 to 63, picks one character from the alphabet.
- “Man” as bytes in binary:
01001101 01100001 01101110 - Joined and regrouped in sixes:
010011 010110 000101 101110 - As numbers: 19, 22, 5, 46
- As alphabet positions: T W F u → TWFu
Because every 3 bytes become 4 characters, Base64 output is about 33% larger than the input.
Padding with =
If the input does not divide evenly into 3-byte groups, the final group is completed with zero bits and the missing characters are written as =. These are the test vectors from RFC 4648 §10, produced by this tool:
| Input | Base64 |
|---|---|
f | Zg== |
fo | Zm8= |
foo | Zm9v |
foob | Zm9vYg== |
fooba | Zm9vYmE= |
foobar | Zm9vYmFy |
Untick Padding to drop the trailing = signs; RFC 4648 allows that when the data length is known some other way (as in many URL and JWT uses). The decoder accepts input with or without padding.
Standard vs URL-safe Base64
+ and / have special meanings in URLs and file paths, so RFC 4648 §5 defines a “base64url” alphabet that uses - and _ instead. Tick URL-safe to encode with it. For example, the two bytes FB FF give +/8= in standard Base64 and -_8 in base64url without padding. When decoding, both alphabets are recognised automatically, so you can paste either kind.
UTF-8 and the btoa() problem
Base64 encodes bytes, not letters, so text has to be turned into bytes first. This tool uses UTF-8, where é is two bytes and an emoji is four — é encodes as w6k=. JavaScript's built-in btoa() only accepts characters up to code 255 and throws an error on emoji, which is a common source of bugs; doing the UTF-8 step explicitly, as here, avoids it. On decoding, the bytes are read as UTF-8, and if they are not valid text (for example a PNG image), you will see a warning.
Tips
- Whitespace and line breaks in pasted Base64 (as in email MIME blocks) are ignored.
- Characters outside the alphabet are reported rather than silently dropped into the output.
- Base64 is not secret. Anyone can decode it; use real encryption for sensitive data.
To see the raw bytes behind any text, use the hex converter or the binary translator.
Where you meet Base64
- Data URLs embed small files directly in HTML or CSS, in the form
data:image/png;base64,…(RFC 2397). - Email attachments are usually sent with MIME’s base64 content-transfer-encoding, which wraps lines at 76 characters (RFC 2045). The decoder ignores those line breaks.
- JSON Web Tokens join three base64url segments with dots, without padding (RFC 7519). Paste one segment at a time to read the header or payload.
- HTTP Basic authentication sends
username:passwordBase64-encoded (RFC 7617) — a reminder that Base64 hides nothing, so such headers must travel over HTTPS.
How long will the output be?
With padding, the output length is 4 × ⌈n ÷ 3⌉ characters for n input bytes. Without padding, the trailing = signs are simply left off:
| Input bytes | 1 | 2 | 3 | 4 | 5 | 6 | 10 | 100 |
|---|---|---|---|---|---|---|---|---|
| Padded | 4 | 4 | 4 | 8 | 8 | 8 | 16 | 136 |
| Unpadded | 2 | 3 | 4 | 6 | 7 | 8 | 14 | 134 |
Remember that the input is measured in UTF-8 bytes, not characters: a single emoji is four bytes and so becomes eight Base64 characters (with padding) on its own.