Data formats

UTF-16 again

The magic of UTF-16 comes from reserving two 1024-element subranges of the 0 to 0x10000 range covered by the BMP to not be used to encode characters.

All the numbers in the range U+D800 to U+DBFF (inclusive) are the high surrogates

All the numbers in the range U+DC00 to U+DFFF (inclusive) are the low surrogates

Recognizing surrogate pairs

All the surrogate pairs have their five most-significant bits in the pattern 0xd8 or 11011.

Thus we can check whether a given 16-bit value, c, is a surrogate via (c & 0xf800) == 0xd800

Encoding a 20-bit value

Given two surrogates (high and low) we can then extract the bottom 10 bits of each by masking with 0x3ff.

Then we shift the ten bits fromm the high surrogate left ten and or with the ten bits from the low surrogate and we have a 20-bit value.

Then add the 0x10000, the number of elements in the BMP.

int cp = (((high & 0x3ff) << 10) | (low & 0x3ff)) + 0x10000;

Endianness

Everything is stored as a sequence of bytes.

So how do we store a 16-bit value?

We write a number like 0xcafe.

But if we want to store that to a file as two bytes, do we write ca then fe or fe and then ca?

Also, how are they stored in memory?

This is a super famous computer science paper.