Text as data

What, even, is a string?

Strings are made up of characters.

Okay, what’s a character?

For our purposes, a character is a letter or digit or punctuation mark. Space is also a character. And newline. Basically the stuff we build text out of.

But computers don’t know about characters. So we need to encode characters as numbers.

In the olden days, a character was just a number from 0-127

ASCII

American Standard Code for Information Interchange

  0 nul    1 soh    2 stx    3 etx    4 eot    5 enq    6 ack    7 bel
  8 bs     9 ht    10 nl    11 vt    12 np    13 cr    14 so    15 si
 16 dle   17 dc1   18 dc2   19 dc3   20 dc4   21 nak   22 syn   23 etb
 24 can   25 em    26 sub   27 esc   28 fs    29 gs    30 rs    31 us
 32 sp    33  !    34  "    35  #    36  $    37  %    38  &    39  '
 40  (    41  )    42  *    43  +    44  ,    45  -    46  .    47  /
 48  0    49  1    50  2    51  3    52  4    53  5    54  6    55  7
 56  8    57  9    58  :    59  ;    60  <    61  =    62  >    63  ?
 64  @    65  A    66  B    67  C    68  D    69  E    70  F    71  G
 72  H    73  I    74  J    75  K    76  L    77  M    78  N    79  O
 80  P    81  Q    82  R    83  S    84  T    85  U    86  V    87  W
 88  X    89  Y    90  Z    91  [    92  \    93  ]    94  ^    95  _
 96  `    97  a    98  b    99  c   100  d   101  e   102  f   103  g
104  h   105  i   106  j   107  k   108  l   109  m   110  n   111  o
112  p   113  q   114  r   115  s   116  t   117  u   118  v   119  w
120  x   121  y   122  z   123  {   124  |   125  }   126  ~   127 del

7 bits. Fits in a byte and encodes 128 characters.

Gave us ASCII art

                        .="=.
                      _/.-.-.\_     _
                     ( ( o o ) )    ))
                      |/  "  \|    //
      .-------.        \'---'/    //
     _|~~ ~~  |_       /`"""`\\  ((
   =(_|_______|_)=    / /_,_\ \\  \\
     |:::::::::|      \_\\_'__/ \  ))
     |:::::::[]|       /`  /`~\  |//
     |o=======.|      /   /    \  /
jgs  `"""""""""`  ,--`,--'\/\    /
                   '-- "--'  '--'

And the rest of the world?

Turns out there’s more than 128 characters in the world. Even just counting languages that use an alphabet.

In many European languages they use characters with diacritics (e.g. accents, cedillas, umlauts).

In Russia, Greece, and elsewhere they use completely other alphabets.

These characters need numbers too.

ISO 8859

Extended ASCII by using the numbers from 128-255.

But kept compatibility with ASCII because US hegemony is a thing.

There are 16 flavors of ISO-8859-1 to ISO-8859-16 (though ISO-8859-12 was abandoned, so 15) that each use the numbers from 128 to 255 to encode different extra characters.

Obviously, to understand data encoded this way, you need to know which of the fifteen flavors is being used.

Asia would like a word

Or maybe a hundred thousand words.

Asian languages such as Chinese, Japanese, and Korean don’t use nice tidy alphabets that can be composed to form words.

Instead they have thousands and thousands of characters.

These characters need numbers too.

Asian encodings

Chinese, Japanese, and Korean computer makers had had to roll their own encoding schemes for mapping their characters to numbers.

They had to wrestle with the trade offs of how to encode a bigger set of numbers efficiently.

They were complicated and there were many flavors.

The Japanese invented the word mojibake to refer to text that has been rendered using the wrong character encoding resulting in gibberish looking characters.

Putting it all together

In the 80s people started trying to make a coded character set that would cover all characters in all languages, mapping unique numbers, called code points, to individual characters in a repertoire of possible characters.

At first they were aiming for a repertoire covering the characters in all modern languages and figured there were no more than \(2^{14}\), so they could be encoded with 16-bit quantities.

Hegemony: still with us

Because most text on computers was still in English and other Western European languages, the first 256 code points in Unicode were the same numbers as ISO-8859-1, the Western European flavor of ISO-8859.

And the first 128 code points of ISO-8859-1 were just the old ASCII code points.

Handy. But still a bit hegemonic.

Java was Unicode from the start

Work on Java started in June 1991. Unicode 1.0 came out in October of 1991 and several more versions were released, up to version 1.1.5, released in July of 1995.

The Unicode people swore that Unicode code points would always fit in 16-bits.

So Java embraced Unicode and defined a 16-bit char type to hold Unicode code points and defined String as a thin wrapper over a char[].

Then Unicode screwed Java

Java 1.0 was released January 1996.

About six months later, in July 1996, the Unicode people released Unicode 2.0.

“LOL, JK. We need 21 bits to represent all Unicode code points.”

A code point for every character

Unicode had decided that they wanted to cover all characters: modern, historical, and future.

So they defined a plane as a set of \(2^{16}\) code points and called the original set of code points the Basic Multilingual Plane or BMP.

Then they added sixteen other planes. U+10000-U+1FM, U+20000-U+2FFFF, etc. up to U+100000-U+10FFFF.

Life in the BMP

And all the commonly used characters in modern languages live in the BMP, i.e. they have code points less than \(2^{16}\) and could be stored in a Java char.

Except emoji. 😭

(Also, historical scripts and some of the more obscure Chinese characters live outside the BMP.)

Character encoding

Code points are just numbers and there are lots of ways to encode them.

In the simpler days of ASCII and ISO-8859, the encoding was trivial: each number fit in a byte so just use a byte.

In Unicode 1.0 you could just use a 16-bit number though that would have the effect of taking twice as much space to store all the characters that fit in 8-bits. Which included almost all of the characters in text in English and other European languages.

Fixed-length encodings

A fixed-length encoding uses the same number of bytes for every code point. Unicode code points need 21 bits which is at least three bytes. But computers don’t really deal with three bytes at a time so we’d use an int and spend four bytes per code point. This is called UTF-32.

The nice thing about fixed-length encodings is that you can figure out offsets using just math. Character \(n\) in a UTF-32-encoded string is the four bytes \(4n\) bytes from the beginning of the string.

The downside

That’d be fine if all code points were used equally often.

But in practice, most text is made up of the same old ASCII code-points we’ve had since the 60s.

UTF-32 wastes a ton of space!

Variable-length encodings

The solution is to encode different code points in different numbers of bytes.

Unicode defines two variable length encodings.

Each encoding has a characteristic code unit which is the data type used in the encoding.

Those code units are then combined to encode code points.

UTF-16

UTF-16 uses 16-bit code units (so basically a Java char) and can encode all code points using either one 16-bit code unit (for all the code points in the BMP) or a pair of 16-bit code units for everything else.

The trick is they reserved 2,048 of the 65,536 16-bit values in the BMP to be used as surrogate pairs where we take ten bits from one 16-bit value, and ten bits from the other and smoosh them together to make a twenty-bit quantity. Then we add 0x10000 to get out of the BMP and into the higher planes.

UTF-8

UTF-8 uses an 8-bit code unit (i.e. a byte) but uses between one and four code units to encode each code point.

Since it encodes the original ASCII code points in just one byte, plain ASCII text and UTF-8-encoded text that contains only ASCII code points is identical.

And it can encode the code points for many non-Latin alphabets such as Greek, Cyrillic, Coptic, Armenian, Hebrew, and Arabic, as well as most diacritical marks in just two bytes.

Today

Almost all text in files and flowing over the internet will be encoded in UTF-8.

Java String can no longer be simply a wrapper over a char[] with each char representing a code point.

Instead, the char[] holds UTF-16 encoded code points so some pairs of chars will encode just one code point.

Consequences

The length() of a String is its length in code units, not code points.

"😭".length() returns 2.

The substring method may not work at some indices.

"😭".substring(0,1) and "😭".substring(1, 2) both return "?" since they each access just one code unit of two-code-unit pair.

Key takeaways

  • Computers don’t know about text
  • We use numbers to represent characters
  • Unicode defines numbers for pretty much every character humans use
  • Numbers need to be encoded into bytes
  • There are different encodings: UTF-16 and UTF-8 are the important ones.
  • History is baked into the format.