Skip to content
IGCSE·Tuition
Computer Science · Lessons

Explain character encoding from a supplied example

A question hands you a code for one letter and expects you to work out the rest, which feels unfair until you see the pattern.

On this page
  1. How do you work out a code from a supplied example?
  2. What is the difference between ASCII and Unicode?
  3. Worked example
  4. The mistake to watch for
  5. Check yourself
  6. Where this leads next

A computer stores text by giving every character a number, called its character code, and storing that number in binary. A character set such as ASCII or Unicode is the agreed table that links characters to codes.

Exam questions usually supply one or two codes and ask you to find others, convert them to binary, or explain why a bigger character set is needed. This lesson belongs to Representing numbers and text and relies on binary and denary conversion.

How do you work out a code from a supplied example?

Letters are stored in alphabetical order, so their codes are consecutive. If you are told that A is 65, then B is 66, C is 67, and so on. To find a letter, count how many places it is after A and add that to 65.

LetterPosition after ACode
A065
B166
C267
G671
O1479

The same idea works for lowercase. If a is 97, then e is 97 + 4 = 101.

What is the difference between ASCII and Unicode?

ASCII uses 7 bits for each character, so it has 27 = 128 different codes. That is enough for English letters, digits, punctuation and some control codes, but not for the characters of other writing systems. Extended versions use 8 bits for 256 codes.

Unicode uses more bits per character, so it can hold characters from many languages, plus symbols and emoji. A student writing 你好 in a message needs Unicode, because ASCII has no codes for those characters. Unicode keeps the same codes as ASCII for the first 128 characters, so older text still works.

The trade-off is storage: more bits per character means a larger file for the same number of characters.

Worked example

You are told that the ASCII code for the capital letter A is 65.

(a) Decode the codes 71 and 79.

71 − 65 = 6, so the letter is 6 places after A. A, B, C, D, E, F, G: that is G.

And 79 − 65 = 14, which is 14 places after A, giving O. The word is GO.

(b) Write the code for C in 8-bit binary.

C is 65 + 2 = 67. Using place values: 67 = 64 + 2 + 1, so the binary is 01000011. Check: 64 + 2 + 1 = 67.

(c) The lowercase letter a has code 97. What is the code for lowercase c, and why is it different from capital C?

Lowercase c is 97 + 2 = 99, which is 01100011 (64 + 32 + 2 + 1 = 99). It differs from C (67) because capital and lowercase are different characters, so each needs its own code. The gap is 97 − 65 = 32.

The mistake to watch for

A common slip is confusing a character with its numeric value.

Question: The character ‘7’ is typed. What code is stored if the digit ‘0’ has code 48?

Mistaken answer: 7 (or 00000111).

The character ‘7’ is a symbol, so it has its own code: 48 + 7 = 55, which is 00110111 (32 + 16 + 4 + 2 + 1 = 55). The number 7 and the character ‘7’ are different things, and a program stores them differently. A good test is to ask whether the 7 will be added up or only displayed as text.

Check yourself

1. Given that a is 97, what is the code for e?

Show answer

e is 4 places after a, so 97 + 4 = 101.

2. Given that A is 65, find the code for J in 7-bit binary.

Show answer

J is the 10th letter, so it is 9 places after A: 65 + 9 = 74. Binary: 74 = 64 + 8 + 2, so 1001010 (7 bits). Check: 64 + 8 + 2 = 74.

3. Why does a messaging app need Unicode rather than 7-bit ASCII?

Show answer

7-bit ASCII has only 128 codes, which is not enough for letters from other languages such as Chinese, Tamil or Arabic, or for emoji. Unicode uses more bits per character, so it can give each of them its own code.

Where this leads next

The next lesson, relating bit depth to possible representations, generalises the “how many codes?” question. Test yourself in the mixed practice set. A teacher can help you build short, accurate explanations for these questions in online one-to-one Computer Science tuition.

Questions people ask

What is a character set?

A character set is an agreed table that gives every character, such as a letter, digit or symbol, its own number. A computer stores that number in binary. Without an agreed table, two computers would not read the same bits as the same text.

What is the difference between ASCII and Unicode?

ASCII uses 7 bits per character, giving 128 characters, which covers English letters, digits and common symbols. Unicode uses more bits per character so it can include letters from many languages, symbols and emoji. Unicode includes the ASCII codes at the start.

Why does a capital letter have a different code from a lowercase one?

They are different characters, so they need different numbers. In ASCII, A is 65 and a is 97, a difference of 32. The same gap of 32 applies to every letter pair, which makes it easy to work out a lowercase code from a capital.

Do I need to memorise the whole ASCII table?

No. Questions normally supply a starting code or a table. Learn the idea that letters are in alphabetical order with consecutive codes, then work out the others from the code you are given. Check your syllabus year for what Cambridge expects.

Sources

  1. Cambridge IGCSE Computer Science 0478 syllabus page

Updated:

Your next step

If character code questions leave you unsure whether to count, convert or explain, a one-to-one teacher can show you which of the three the question wants and practise the wording with you.

Paid one-hour trial at your assigned teacher’s confirmed rate, starting from RM80. Other fees, schedules and ongoing arrangements are confirmed directly with your teacher after the trial class.

Tuition is arranged with a parent or guardian. Send them this page on WhatsApp and they can enquire for you.

Parent or guardian? Enquire here

9,000+ students helped through our service