A computer stores text by giving every character a number, called its character code, and storing that number in binary. A character set such as ASCII or Unicode is the agreed table that links characters to codes.
Exam questions usually supply one or two codes and ask you to find others, convert them to binary, or explain why a bigger character set is needed. This lesson belongs to Representing numbers and text and relies on binary and denary conversion.
How do you work out a code from a supplied example?
Letters are stored in alphabetical order, so their codes are consecutive. If you are told that A is 65, then B is 66, C is 67, and so on. To find a letter, count how many places it is after A and add that to 65.
| Letter | Position after A | Code |
|---|---|---|
| A | 0 | 65 |
| B | 1 | 66 |
| C | 2 | 67 |
| G | 6 | 71 |
| O | 14 | 79 |
The same idea works for lowercase. If a is 97, then e is 97 + 4 = 101.
What is the difference between ASCII and Unicode?
ASCII uses 7 bits for each character, so it has 27 = 128 different codes. That is enough for English letters, digits, punctuation and some control codes, but not for the characters of other writing systems. Extended versions use 8 bits for 256 codes.
Unicode uses more bits per character, so it can hold characters from many languages, plus symbols and emoji. A student writing 你好 in a message needs Unicode, because ASCII has no codes for those characters. Unicode keeps the same codes as ASCII for the first 128 characters, so older text still works.
The trade-off is storage: more bits per character means a larger file for the same number of characters.
Worked example
You are told that the ASCII code for the capital letter A is 65.
(a) Decode the codes 71 and 79.
71 − 65 = 6, so the letter is 6 places after A. A, B, C, D, E, F, G: that is G.
And 79 − 65 = 14, which is 14 places after A, giving O. The word is GO.
(b) Write the code for C in 8-bit binary.
C is 65 + 2 = 67. Using place values: 67 = 64 + 2 + 1, so the binary is 01000011. Check: 64 + 2 + 1 = 67.
(c) The lowercase letter a has code 97. What is the code for lowercase c, and why is it different from capital C?
Lowercase c is 97 + 2 = 99, which is 01100011 (64 + 32 + 2 + 1 = 99). It differs from C (67) because capital and lowercase are different characters, so each needs its own code. The gap is 97 − 65 = 32.
The mistake to watch for
A common slip is confusing a character with its numeric value.
Question: The character ‘7’ is typed. What code is stored if the digit ‘0’ has code 48?
Mistaken answer: 7 (or 00000111).
The character ‘7’ is a symbol, so it has its own code: 48 + 7 = 55, which is 00110111 (32 + 16 + 4 + 2 + 1 = 55). The number 7 and the character ‘7’ are different things, and a program stores them differently. A good test is to ask whether the 7 will be added up or only displayed as text.
Check yourself
1. Given that a is 97, what is the code for e?
Show answer
e is 4 places after a, so 97 + 4 = 101.
2. Given that A is 65, find the code for J in 7-bit binary.
Show answer
J is the 10th letter, so it is 9 places after A: 65 + 9 = 74. Binary: 74 = 64 + 8 + 2, so 1001010 (7 bits). Check: 64 + 8 + 2 = 74.
3. Why does a messaging app need Unicode rather than 7-bit ASCII?
Show answer
7-bit ASCII has only 128 codes, which is not enough for letters from other languages such as Chinese, Tamil or Arabic, or for emoji. Unicode uses more bits per character, so it can give each of them its own code.
Where this leads next
The next lesson, relating bit depth to possible representations, generalises the “how many codes?” question. Test yourself in the mixed practice set. A teacher can help you build short, accurate explanations for these questions in online one-to-one Computer Science tuition.