How the binary translator works
Computers store text as numbers, and numbers as bits. This tool turns each character into its UTF-8 bytes and writes every byte as eight binary digits. For the basic Latin letters, digits and punctuation, UTF-8 is identical to ASCII, so one character is one byte. Characters beyond that — accented letters, symbols such as €, and emoji — take two, three or four bytes, exactly as defined in RFC 3629. Decoding reverses the process and reassembles multi-byte characters.
The Binary box is tolerant. You can paste bytes separated by spaces, commas or new lines, a single unbroken string of bits, or values with a 0b prefix. If the groups are separated and some are shorter than eight bits (common with 7-bit ASCII exercises), each group is read as one byte: 1001000 1101001 reads as “Hi”. Anything that is not a 0 or 1 is ignored and reported, and if the bit count is not a multiple of eight the leftover bits are flagged instead of silently guessed.
Worked example: “Hi” in binary
| Character | Decimal code | Binary byte |
|---|---|---|
| H | 72 | 01001000 |
| i | 105 | 01101001 |
Each bit position is worth double the one to its right. Reading the byte for H, 01001000, from left to right:
| 128 | 64 | 32 | 16 | 8 | 4 | 2 | 1 |
|---|---|---|---|---|---|---|---|
0 | 1 | 0 | 0 | 1 | 0 | 0 | 0 |
The 1s sit under 64 and 8, which add up to 72 — the code for capital H. The same method works for any byte: add up the place values wherever there is a 1.
One character, up to four bytes
UTF-8 marks how many bytes a character uses with the leading bits of the first byte: 0xxxxxxx for one byte, 110xxxxx for two, 1110xxxx for three and 11110xxx for four. Every following byte starts with 10. That is why the bytes below look so regular:
| Character | Code point | Bytes | UTF-8 binary |
|---|---|---|---|
| A | U+0041 | 1 | 01000001 |
| é | U+00E9 | 2 | 11000011 10101001 |
| € | U+20AC | 3 | 11100010 10000010 10101100 |
| 😀 | U+1F600 | 4 | 11110000 10011111 10011000 10000000 |
If you decode binary that was made with a different encoding (for example Latin-1, where é is the single byte 11101001), the bytes will not form valid UTF-8. The translator shows such bytes as � and says so, rather than producing the wrong letter.
Bits, bytes and spaces
A bit is one binary digit; a byte is eight bits, which can hold 256 different values (0 to 255). ASCII itself only needs seven bits (0–127), which is why you sometimes see seven-digit groups in textbooks; padding them to eight with a leading 0 gives the same value. Spaces between bytes are only for readability — they are not part of the data — so you can switch the separator to “None” for a compact string, or “New line” to list one byte per line.
Related tools and charts
To look up a single letter, use the printable binary alphabet chart. For other number bases try the hex to text converter (two hex digits per byte) or the ASCII converter (decimal codes), and for binary data in text form, the Base64 encoder.
Reading binary by hand, step by step
- Remove spaces and split the bits into groups of eight from the left.
- Convert each group with the place values 128, 64, 32, 16, 8, 4, 2, 1.
- Look each number up in an ASCII table. If a byte is 192 or more, it starts a multi-byte UTF-8 character — combine it with the following bytes that start with
10.
Common mistakes
- A dropped digit. One missing bit shifts every following byte, turning the rest of the message into nonsense. The translator flags a bit count that is not a multiple of eight.
- Mixing 7-bit and 8-bit groups. Separate the groups with spaces and the translator reads each group as one byte, whatever its length.
- Reading the bits as one big number.
01001000 01101001is two separate bytes (72 and 105), not the single number 18,537.