{"article":{"slug":"utf-8000-unlimited-utf-8","title":"UTF-8000: Unlimited UTF-8","subtitle":null,"summary":"A playful but carefully argued encoding proposal that extends UTF-8 to arbitrary length while preserving ASCII compatibility and self-synchronization, with a reference pipx implementation.","content_type":"essay","language":"en","canonical_url":"https://utf-8000.jb2170.com/","author":{"name":"jb2170","url":"https://jb2170.com/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"jb2170","url":"https://utf-8000.jb2170.com/","listing_slug":null,"listing":null},"topics":[{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Unicode","slug":"unicode","url":"https://listedarticles.com/topics/unicode"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"},{"name":"Encoding","slug":"encoding","url":"https://listedarticles.com/topics/encoding"},{"name":"Systems Programming","slug":"systems-programming","url":"https://listedarticles.com/topics/systems-programming"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":20405,"reading_minutes":89,"published_at":"2026-07-04T00:00:00.000Z","added_at":"2026-09-20T09:06:37.385Z","updated_at":"2026-09-20T09:06:37.385Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/utf-8000-unlimited-utf-8","markdown_url":"https://listedarticles.com/articles/utf-8000-unlimited-utf-8.md","example":false,"citation":"jb2170, jb2170. \"UTF-8000: Unlimited UTF-8.\" 4 Jul 2026. https://utf-8000.jb2170.com/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://utf-8000.jb2170.com/"},"body_markdown":"# UTF-8000\n\nUnlimited UTF-8! ASCII ⊆ UTF-8 ⊆ UTF-8000.\n\nNo special cases introduced. All properties preserved.\n\nTry out the reference implementation with `$ pipx install UTF-8000`.\n\n                UTF-8000 is in no way endorsed by or representative of the Unicode Consortium. \n\n                This is a fun standalone project / proposal.\n              \n\n## TLDR / Examples\n\n| ASCII |  |  |  |  |  |  |  |  |  |  |  | \n| 1 | `0xxxxxxx` |  |  |  |  |  |  |  |  |  |  | \n| UTF-8 |  |  |  |  |  |  |  |  |  |  |  | \n| 2 | `110xxxxx` | `10xxxxxx` |  |  |  |  |  |  |  |  |  | \n| 3 | `1110xxxx` | `10xxxxxx` | `10xxxxxx` |  |  |  |  |  |  |  |  | \n| 4 | `11110xxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  |  |  |  |  |  |  | \n| UTF-8000 |  |  |  |  |  |  |  |  |  |  |  | \n| 5 | `111110xx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  |  |  |  |  |  | \n| 6 | `1111110x` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  |  |  |  |  | \n| 7 | `11111110` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  |  |  |  | \n| 8 | `11111111` | `100xxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  |  |  | \n| 9 | `11111111` | `1010xxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  |  | \n| 10 | `11111111` | `10110xxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  | \n| ... |  |  |  |  |  |  |  |  |  |  |  | \n| 22 | `11111111` | `10111111` | `10111111` | `10110xxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` | \n| ... |  |  |  |  |  |  |  |  |  |  |  | \n\nThere is nothing special-case-y about the example 22-byte code unit here. It is just a good prototypical example, demonstrating the power of UTF-8000 with multiple start bytes.\n\nThere are only two special cases, both of which are inherited from UTF-8: ASCII as is, and 2-byte UTF-8 having 4 mandatory content bits to check against overlong encoding as opposed to 5 for all longer length code units.\n\n## Anatomy\n\nHere is anatomical diagram of the example 22-byte code unit from the tldr.\n\nSee the glossary for more information on the definitions of the terms.\n\nByte number four is exciting! It is a continuation byte, a start byte, the final start byte, has content bits, and has only some of the mandatory content bits, which are straddled across the final start byte and first non-start byte.\n\nThe main contribution of UTF-8000's specification is clarity on splitting the highest bits of the first byte of UTF-8 code units into self-synchronization bits and start bits, and then making it clear how to stripe the start bits across the continuation bytes if needed, to achieve arbitrarily large code units.\n\n## Glossary\n\nThese terms are ordered somewhat by chronology of first requirement, rather than alphabetically, for convenience.\n\nTerms used within definitions are underlined clickable hyperlinks.\n\n| Term | Definition | \n|---|---|\n| Codepoint | A non-negative integer, aka an unsigned integer. | \n| Code Unit | A sequence of UTF-8000 bytes that encode a single codepoint. | \n| First Byte | The first, one and only, byte that begins a UTF-8000 code unit.                          The self-synchronization prefix of a first byte is either                         `0` for ASCII or`11` for multi-byte code units. This term is **not** synonymous with start byte. A first byte is necessarily a start byte, but not the other way around. It is for this reason that first byte is sometimes also known as first start byte.                          Fun observation: because of the self-synchronization prefix                         `0` the upper hex nibble of ASCII bytes can only be one of`0` ,`1` ,`2` ,`3` ,`4` ,`5` ,`6` ,`7` . This term is mutually exclusive with continuation byte due to self-synchronization. | \n| Continuation Byte | A byte beyond the first byte of a multi-byte UTF-8000 code unit.                          The self-synchronization prefix of a continuation byte is `10` , which is also known as the continuation prefix bits.                          Fun observation: because of the self-synchronization prefix                         `10` the upper hex nibble of continuation bytes can only be one of`8` ,`9` ,`A` ,`B` . This term is mutually exclusive with first byte due to self-synchronization. | \n| Self-Synchronization Prefix | The highest bits of every UTF-8000 byte that indicate whether it is a first byte or a continuation byte. The possible self-synchronization prefixes form a prefix-free tree: ```    .---- ``` `0` First byte for ASCII   `----`1` ---`0` Continuation byte for multi-byte UTF-8000         `----`1` First byte for multi-byte UTF-8000This piece of the clever architecture of UTF-8, which UTF-8000 inherits, provides the property of self-synchronization at a byte level: we can instantaneously tell what kind of byte we are looking at, and where it should belong in a code unit, just by looking at these highest bits.                          This is most useful when decoding part of a file encoded in UTF-8000. If we randomly seek through the file to an arbitrary byte, we can unambiguously tell whether we are at a first byte whence we can begin decoding a new code unit immediately, or that we are at a continuation byte whence we need to seek a little further on in order to find the next first byte in order to begin decoding. Nor do we have to process any bytes prior to our seek position in order to discover some global state or the context of the byte we have seek-ed to; a first byte is always unambiguously a first byte wherever it appears, which we can deduce by its self-synchronization prefix being either `0` or`11` . This is useful not only for random access, but also for error recovery. Suppose that we are decoding an error-prone stream of UTF-8000 bytes and that whenever when we encounter an error (e.g. a rogue 0xC0 byte) we wish to `U+FFFD` See the Wikipedia article for self-synchronizing code for more general info. These bits are highlighted in bright cyan. | \n| Start Byte | A byte containing one or more start bits. The start bytes exist contiguously at the beginning of a UTF-8000 code unit. The power of UTF-8000 is that we can have multiple start bytes, to achieve arbitrary code unit lengths, to encode arbitrarily large codepoints. Sometimes it is sensible to colloquially also include ASCII as a start byte when we are talking about the bytes towards the start of a code unit, even though ASCII bytes have no start bits. Every non-ASCII code unit has at least one start byte. The first start byte is the first byte, and it is followed by zero or more continuation bytes that are also start bytes. Therefore because a UTF-8000 code unit can have multiple start bytes, this term is **not** synonymous with first byte.                          In restricting to only UTF-8 without UTF-8000, this term                         *is* synonymous with first byte. This is because UTF-8-length code units only require one start byte, whether using up to 4 bytes in the current UTF-8 standard (RFC 3629 (2003)), or using up to 6 bytes in former standards (RFC 2044                         (1996) and                         RFC 2279                         (1998)). | \n| Start Bits | The unary-code sequence of bits contained in the start bytes of a multi-byte UTF-8000 code unit that tells us the length of the code unit in bytes.                          For a code unit made of `n` bytes the start bits are`n-2``1` bits followed by a terminating`0` bit. To be clear, the start bits include this terminating zero bit. Thus the start bits sequence is of length`n-1` and looks like`111...10` . The possible start bits sequences form a prefix-free tree: ```    .---- ``` `0` Two byte UTF-8   `----`1` ---`0` Three byte UTF-8         `----`1` ---`0` Four byte UTF-8               `----`1` ---`0` Five byte UTF-8000                     `----...    n byte UTF-8000For an `n` byte code unit where`n < 8` the start bits all fit together snugly in the first byte. Otherwise they are striped across as many of the first few bytes as they need, filling the free bits that are not occupied by continuation prefix bits.                          This is another piece of the clever architecture of UTF-8, which UTF-8000 inherits, that provides the property of self-punctuation also known as a `0` bit, we know exactly how many bytes we expect in that code unit. Notwithstanding errors we can therefore succeed in decoding the code unit by reading exactly that many bytes, and no more. This avoids a problem of dumber variable-length encodings whose code units do not intrinsically indicate their length: one has to read beyond the last byte of a code unit, that is one reads the first byte of the next code unit, in order to know that the current code unit has finished. For very dumb encodings which have neither self-synchronization nor self-punctuation, to make random access possible one would have to put dedicated auxiliary bytes,  See the Wikipedia articles for prefix code and unary coding for more general info. This term is mutually exclusive with content bits. These bits are highlighted in bright magenta. | \n| Content Byte | A byte containing one or more content bits.                          A byte being a content byte does not imply that it is a continuation byte. For example a 3-byte code unit begins with `1110xxxx` , which contains 4 content bits and is not a continuation byte.                          A byte being a continuation byte does not imply that it is a content byte. For example a 22-byte code unit contains                         `10111111` as its second byte, which is a continuation byte and has no content bits. | \n| Content Bits |                          The sequence of bits in a code unit beyond the start bits and to the end of the code unit, in which the codepoint's binary bits are stored. For example a 3-byte code unit, which has the form                         `1110xxxx``10xxxxxx``10xxxxxx` , has 16 content bits.                          For ASCII there are 7 content bits. These seven bits `xxxxxxx` combined with a byte's highest bit being set to the self-synchronization prefix`0` means that ASCII is perfectly included into UTF-8 without being altered. Thus ASCII code units take the form`0xxxxxxx` . Otherwise for an `n` byte code unit, where`n > 1` , there are`5n+1` content bits. This is how we arrive at that formula: We start with`n` blank bytes, each of which has`8` bits. For each byte`2` bits are taken by the self-synchronization prefix. Then an additional`n-1` bits are taken by the start bits. Thus there are`8n - 2n - (n-1) = 5n+1` bits left for content bits. Another way to think about the`5` in this formula is by extending from`n-1` bytes to`n` bytes by appending another continuation byte. By doing this we gain`6` free bits in the continuation byte, but we lose`1` bit to the longer start bits sequence, thus overall we gain`6-1 = 5` bits for content bits. This term is mutually exclusive with start bits. These bits are highlighted in lime. | \n| Mandatory Content Byte | A byte containing one or more mandatory content bits. These are the bytes we check for overlong encoding when decoding a code unit. | \n| Mandatory Content Bits |                          The first 0, 4, or 5 content bits of a code unit in which there must be at least one `1` bit, lest the bytes form an overlong encoding, which is forbidden. For ASCII there are 0 mandatory content bits, and thus no anti-overlong checking is required. This is because ASCII is the smallest possible code unit. For 2-byte UTF-8000 there are 4 mandatory content bits. This is because in the jump from 1-byte ASCII to 2-byte UTF-8 we jump from 7 content bits to 11 content bits. Thus the number of content bits we gain is 11 minus 7 which is 4. Otherwise for `n` byte UTF-8000, where`n > 2` , there are 5 mandatory content bits. This is because in the jump from`n-1` byte UTF-8000 to`n` byte UTF-8000 we add on an extra continuation byte, which has 6 free bits, but we lose 1 bit to the longer start bits sequence. Thus overall the number of content bits we gain is 6 minus 1 which is 5. Read about overlong encoding for why mandatory content bits are of interest. These bits are highlighted in bright lime. | \n| Overlong Encoding | Forbidden                           For example one could incorrectly try to encode the codepoint 0x41, 65, ASCII capital A, using 2-byte UTF-8 as                         `11000001``10000001` . Observe that all the mandatory content bits are`0` which is the definition an overlong encoding. This indicates that we could have encoded 0x41 in a shorter code unit, in this case as ASCII`01000001` .                          Security is one main reason why we forbid overlong encoding. For example we ensure that                         `11100000``10000000``10000000` cannot be decoded as codepoint 0, the null`strcpy(3)` and friends.`strcpy` would not interpret this Uniqueness of encoding is another reason why we forbid overlong encoding. Every codepoint has one unique valid representation as a UTF-8000 code unit, which is easy to encode and decode using bitshifting.                          Fun observation: because all 4 of 2-byte UTF-8's mandatory content bits lie in the first-and-final start byte, we can explicitly rule out                         `11000000` (0xC0) and`11000001` (0xC1) as permanently invalid bytes. They will never ever appear anywhere in a valid UTF-8000                         code unit! | \n\n## Properties\n\nMany of these properties of UTF-8000 are explained in detail in an appropriate section of the glossary and hyperlinks to the glossary are provided.\n\n### Bit Counts\n\nThe number of content bits and mandatory content bits are very predictable as a function of `n`, the length of a code unit.\n\n| code unit length | number of content bits | number of mandatory content bits | \n|---|---|---|\n| `n = 1` | `7` | `0` | \n| `n = 2` | `5n+1 ( = 11)` | `4` | \n| `n > 2` | `5n+1` | `5` | \n\n### Why the Special Cases?\n\nAs stated in the tldr, there are only two special cases, both of which are inherited from UTF-8:\n\n**1-byte UTF-8 (ASCII)** which has two points of interest:\n\n- It has 7 content bits which does not fit the pattern of `5n+1` . See the glossary section for content bits for an explanation, and see the rejected alternative ASCVI code for a version of UTF-8 if ASCII were 6 bit instead of 7 bit which eliminates this special case.\n- ASCII has 0 mandatory content bits because it cannot possibly be overlong since it is the smallest possible code unit. This is fine.\n\n**2-byte UTF-8** which has one point of interest:\n\n- It has 4 mandatory content bits, as opposed to 5 for all longer code units. See the glossary section for mandatory content bits for an explanation.\n\nThe remarkable fact that UTF-8000 does not introduce any new special cases in extending UTF-8 is confirmation to me that this is the canonical, correct way to extend UTF-8. In other words UTF-8 in its current restricted 4 byte form *is* UTF-8000, but only a small part of it.\n\n                The fact that we are even able to extend in the first place is also testament to the clever planning and care that Ken Thompson and Rob Pike put into the architecture of UTF-8, which we ensure to maintain as we extend to UTF-8000. Unary code codewords for the start bits sequences, which form a self-similar tree, were a great choice being simple and extensible. In the earliest draft of UTF-8, the six-byte start-byte looked like `111111xx`. This was changed a few days later to `1111110x`. That way the number of content bits is not a special case, and the start bits don't saturate the unary code binary tree, leaving the door open for our future expansion.\n              \n\nThis is why I think of UTF-8 as the capstone of the Unix Philosophy.\n\n### Information Rate\n\nWhat proportion of a code unit is content bits?\n\nFor ASCII this is `7/8 = 87.5%`.\n\nOtherwise for an `n` byte code unit this is `(5n+1) / 8n`, that is `5n+1` content bits out of a total of `8n` bits from `n` bytes. We can rewrite this as `(5/8) + 1/(8n)` which moderately quickly approaches `5/8 = 62.5%`. It is nice that this limit is nonzero and does not depend on `n`.\n\n### Self-Synchronization\n\nInherited from UTF-8 and maintained in UTF-8000.\n\nSee the glossary section for self-synchronization prefix for an explanation of self-synchronization.\n\n                Here's a bit of history: Self-synchronization is one of the reasons why Ken Thompson and Rob Pike decided to design UTF-8, to supersede the earlier FSS-UTF draft by Dave Prosser et al. FSS-UTF proposed a design like eg\n                `110xxxxx`\n                `1xxxxxxx`\n                `1xxxxxxx`\n                for three-byte code units. The problem with it is that one cannot distinguish between first bytes (`110xxxxx`) and continuation bytes (`110xxxxx`) without knowing the prior history of a stream. The UTF-8 fix is to make first byte and continuation byte values disjoint from each other, as one can witness in the byte map below. I have not put Prosser's draft into the rejected ideas section as it has already been formally addressed and superseded by UTF-8.\n              \n\n### Self-Punctuation\n\nInherited from UTF-8 and maintained in UTF-8000.\n\nSee the glossary section for start bits for an explanation of self-punctuation.\n\n### Byte Map\n\nExtended from UTF-8, making use of the higher value bytes. Based off Wikipedia's UTF-8 Byte Map.\n\n|  | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | A | B | C | D | E | F | \n|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|\n| 0 | ␀ | ␁ | ␂ | ␃ | ␄ | ␅ | ␆ | ␇ | ␈ | ␉ | ␊ | ␋ | ␌ | ␍ | ␎ | ␏ | \n| 1 | ␐ | ␑ | ␒ | ␓ | ␔ | ␕ | ␖ | ␗ | ␘ | ␙ | ␚ | ␛ | ␜ | ␝ | ␞ | ␟ | \n| 2 | ␠ | ! | \" | # | $ | % | & | ' | ( | ) | * | + | , | - | . | / | \n| 3 | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | : | ; | < | = | > | ? | \n| 4 | @ | A | B | C | D | E | F | G | H | I | J | K | L | M | N | O | \n| 5 | P | Q | R | S | T | U | V | W | X | Y | Z | [ | \\ | ] | ^ | _ | \n| 6 | ` | a | b | c | d | e | f | g | h | i | j | k | l | m | n | o | \n| 7 | p | q | r | s | t | u | v | w | x | y | z | { | \\| | } | ~ | ␡ | \n| 8 |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  | \n| 9 |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  | \n| A |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  | \n| B |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  | \n| C |  |  | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | \n| D | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | \n| E | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | \n| F | 4 | 4 | 4 | 4 | 4 | 4 | 4 | 4 | 5 | 5 | 5 | 5 | 6 | 6 | 7 | 8+ | \n\nAll bytes except 0xC0 and 0xC1, colored in tomato red, can appear in a valid UTF-8000 stream. See the glossary section for overlong encoding for an explanation of why those two bytes never appear.\n\nASCII, colored in gold yellow, occupies the first half of the table, being 7-bit. Continuation bytes occupy the region colored in sandybrown orange. All other bytes are first bytes for multi-byte code units, whose lengths are indicated in the table.\n\n### `strcmp(3)` Ordering\n\n              Inherited from UTF-8 and maintained in UTF-8000.\n\n                The self-synchronization prefixes of first bytes are monotonically-increasing-ly ordered with respect to code unit length. In other words ASCII is of length 1 and multi-byte is of length greater than 1, and `0` < `11` occupying the highest bits of UTF-8000 bytes.\n              \n\n                The start bit sequences are also monotonically-increasing-ly ordered with respect to code unit length. In other words if `n < m` then `111...[n]...10` `<` `111...[m]...10` as an integer value, occupying the heads of the code unit bytes beyond the self-synchronization prefixes. This would not have been the case had UTF-8 been designed to use the alternative form of unary codewords given by `000...01`.\n              \n\nThe content bits of code units are also monotonically-increasing-ly ordered with respect to codepoint value.\n\nCombining these three things together means that `strcmp(3)`, the C stdlib string comparing function, works the same way on UTF-8000 bytes as it does on UTF-8, as it does on ASCII, effectively comparing the encoded codepoint values against each other without having to actually decode the code units. Nice!\n\n### No Endianness\n\nThe quantum\n\n of ASCII, UTF-8, and UTF-8000 is a single byte. This makes life a breeze! There is no need for a concept of endianness for UTF-8000.\n\nUTF-16 however has a quantum of two bytes, 16-bit units. When writing the codewords in bytes, 8-bit units, should the byte containing the most significant digits or least significant digits be written first? Big-endian, or little-endian? This choice gives UTF-16 two variants, UTF-16-BE and UTF-16-LE. If one cannot predetermine the endianness of a stream, one may wish to use a BOM which is discussed below.\n\n### BOM Support\n\nA Byte Order Mark (BOM) is used at the start of an encoded text stream to indicate what encoding is used. I have never actively used BOMs myself so I've only put a bit of thought into this section.\n\nAs far as I'm aware we don't break BOM support for UTF-8, though we may wish to have a different BOM to strictly distinguish UTF-8 from UTF-8000. Maybe UTF-8000 could have multiple BOMs, one for each integer N greater than or equal to four, to indicate to a decoder the maximum expected code unit length.\n\n                One of the reasons why `U+FFFE` is not a valid Unicode Scalar Value is because `0xFE 0xFF` is the BOM for UTF-16. Since UTF-16 code units are two bytes wide, one may read either `0xFE 0xFF` or `0xFF 0xFE` depending on endianness. To make it clear that `0xFF 0xFE` implies correct for endianness\n\n and cannot be mistaken for a legitimate character, `U+FFFE` is designated as `<noncharacter-FFFE>`. We are relieved in that neither `11111111` `11111110` nor `11111110` `11111111` are valid UTF-8000 sequence extracts, ie UTF-8000 does not introduce incompatibilities with UTF-16.\n              \n\n### Arbitrary Lengths, Sensible Limits\n\nI think we've made it clear by now that UTF-8000 code units can be arbitrarily large. In practice however one *may* wish to set a sensible limit on code unit lengths when decoding. Here we'll discuss a method of finding some nice code unit lengths whose code units store `5n+1 = 2^N` bits, as we are often interested in powers of 2 in computer science.\n\nIt is a common observation that 3-byte UTF-8 stores `5 * 3 + 1 = 16` bits, meaning the Basic Multilingual Plane of Unicode can be encoded in one two and three byte UTF-8. We see that `2^4 mod5 = 16 mod5 = 1 mod5`; if `5n+1` is to be `2^N` for some `n` then certainly `2^N = 1 mod5`. If we enumerate powers of two modulo five then there is a very predictable repeating pattern of 1, 2, 4, 3\n\n. Formally you might say that 2 is a generator of 𝔽<sub>5</sub><sup>*</sup> if you want impress a mathematician! The takeaway is that when `N=4K` for `K≥1` we can find a corresponding `n` such that `5n+1 = 2^N`. We can rewrite `2^N` as `2^(4K) = 16^K`.\n\nIn other words any power of 16 has a UTF-8000 code unit length containing that many bits\n\n. Here are a few of these for reference.\n\n| `K` | `N=4K` | number of content bits `= 2^N` | code unit length `= (2^N - 1) / 5` | \n|---|---|---|---|\n| `1` | `4` | `16` | `3` | \n| `2` | `8` | `256` | `51` | \n| `3` | `12` | `4096` | `819` | \n| `4` | `16` | `65536` | `13107` | \n| `...` | `...` |  |  | \n\nDo remember that strictly speaking one shouldn't allow overlong encodings, if one were for example thinking of storing a small `uint256_t` key in a 51 byte code unit! UTF-8000's variable width nature helps out leading to smaller code units for smaller integers.\n\n## Intuitive Derivation\n\nThere are a few ways that one could arrive at the design for UTF-8000 and the bit counts above.\n\nOne may think to start with UTF-8, notice that the start byte of an `n` byte code unit is prefixed with the unary codeword of length `n+1`, that is `n` `1` bits followed by a `0`, and then figure out how to extend those bits and roll them over into the continuation bytes without losing any important properties. This is what I *originally* did.\n\nWriting this document over a couple of weeks made me introspect the code unit anatomy further, whence I figured out that separating the leading bits into a self-synchronization part and self-punctuation part further illuminates and simplifies the thought process. We shall thus proceed with this perspective.\n\n#### Blank Slate\n\nWe set out to derive the design of an `n` byte code unit, starting out with `n` blank bytes, all of whose bits could possibly be content bits.\n\n                `00000000`\n                `00000000`\n                `00000000`\n                `...`\n                `00000000`\n              \n\n                To achieve self-synchronization we need to distinguish the first byte of the code unit from the continuation bytes that follow. We could do that by setting the highest bit of first bytes to a\n                `0` and to `1` for continuation bytes. Doing it this way round maintains compatibility with ASCII's highest bit being `0`.\n              \n\n                `00000000`\n                `10000000`\n                `10000000`\n                `...`\n                `10000000`\n              \n\nWith the design so far, all code units begin with an ASCII byte. When decoding a code unit, we have no idea whether this first byte actually is ASCII, or it is the first byte of a multi-byte code unit. We want self-punctuation, where a code unit intrinsically tells us how long it is.\n\nTo achieve self-punctuation we create a prefix-free code binary tree, whose leaf node codewords correspond to code unit lengths. These are the start bits sequences. The codeword for `n` shall be embedded inside the code unit towards the start. It must therefore be short enough to fit into the `n` bytes, and reasonably computationally predictable. We try:\n\n```\n  .----\n```\n`0`                          One byte UTF-8 (ASCII)\n  `----`1`---`0`                    Two byte UTF-8\n        `----`1`---`0`            Three byte UTF-8\n              `----`1`---`0`       Four byte UTF-8\n                    `----`1`---`0` Five byte UTF-8000\n                          `----...    n byte UTF-8000\n              This seems reasonably simple so far. We stripe the start bits into the available bits not taken by the self-synchronization prefix. All other bits shall be content bits.\n\n| 1 | `00xxxxxx` |  |  |  |  |  | \n| 2 | `010xxxxx` | `1xxxxxxx` |  |  |  |  | \n| 3 | `0110xxxx` | `1xxxxxxx` | `1xxxxxxx` |  |  |  | \n| ... |  |  |  |  |  |  | \n| 17 | `01111111` | `11111111` | `1110xxxx` | `1xxxxxxx` | ... | `1xxxxxxx` | \n| ... |  |  |  |  |  |  | \n\n                But wait we've broken the distinction of ASCII! We cannot tell the difference between eg\n                `0110xxxx`\n                and\n                `0110xxxx`, or\n                `01111111`\n                and\n                `01111111`. This code would only work if ASCII were six-bit instead of seven-bit. Out of curiosity we investigate this code in the rejected alternatives section ASCVI\n\n.\n              \n\n                To maintain compatibility with ASCII we must treat it as a special case, whereby the self-synchronization prefix `0` is alone sufficient to characterize ASCII. This highlights that the architecting of UTF-8 was not purely a mathematics problem, but was also an engineering problem, working around what already exists.\n              \n\n                Seeing the ASCII-characterizing prefix `0` and the erstwhile continuation prefix `1` as forming a prefix-free tree, albeit only of size two, we must repurpose the the latter codeword as the beginning of the self-synchronization prefixes for first bytes and continuation bytes of multi-byte code units. We choose our new self-synchronization prefixes as `11` for start bytes and `10` for continuation bytes. This produces the following tree:\n              \n\n```\n  .----\n```\n`0`              First byte for ASCII\n  `----`1`---`0` Continuation byte for multi-byte UTF-8000\n        `----`1`        First byte for multi-byte UTF-8000\n              Accordingly adjusting the self-punctuation codewords to apply only to multi-byte code units produces the following tree:\n\n```\n  .----\n```\n`0`                    Two byte UTF-8\n  `----`1`---`0`            Three byte UTF-8\n        `----`1`---`0`       Four byte UTF-8\n              `----`1`---`0` Five byte UTF-8000\n                    `----...    n byte UTF-8000\n              Putting these mechanisms together yields UTF-8000 and we're done!\n\n| ASCII |  |  |  |  |  |  |  |  |  |  |  | \n| 1 | `0xxxxxxx` |  |  |  |  |  |  |  |  |  |  | \n| UTF-8 |  |  |  |  |  |  |  |  |  |  |  | \n| 2 | `110xxxxx` | `10xxxxxx` |  |  |  |  |  |  |  |  |  | \n| 3 | `1110xxxx` | `10xxxxxx` | `10xxxxxx` |  |  |  |  |  |  |  |  | \n| 4 | `11110xxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  |  |  |  |  |  |  | \n| UTF-8000 |  |  |  |  |  |  |  |  |  |  |  | \n| 5 | `111110xx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  |  |  |  |  |  | \n| 6 | `1111110x` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  |  |  |  |  | \n| 7 | `11111110` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  |  |  |  | \n| 8 | `11111111` | `100xxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  |  |  | \n| 9 | `11111111` | `1010xxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  |  | \n| 10 | `11111111` | `10110xxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  | \n| ... |  |  |  |  |  |  |  |  |  |  |  | \n| 22 | `11111111` | `10111111` | `10111111` | `10110xxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` | \n| ... |  |  |  |  |  |  |  |  |  |  |  | \n\nIt is a trivial* result of coding theory that the product of two prefix-free codes is also a prefix-free code. The product of our trees looks like:\n\n```\n  .----\n```\n`0`                                                  ASCII byte\n  `----`1`---`0`                               UTF-8 continuation byte\n        `----`1`---`0`                    Two byte UTF-8    start byte\n              `----`1`---`0`            Three byte UTF-8    start byte\n                    `----`1`---`0`       Four byte UTF-8    start byte\n                          `----`1`---`0` Five byte UTF-8000 start byte\n                                `----...    n byte UTF-8000 start byte\n              Without the color highlighting this is how most people think about UTF-8: a start byte whose prefix of `n` `1` bits and a terminating `0` bit provides both self-synchronization and self-punctuation, and continuation bytes using a short prefix of `10` for space-efficient encoding. This makes sense for a small number of bytes, but the trick to unlock a perspective of infinite extensibility is to split this tree into the self-synchronization part and self-punctuation part; ie we un-product those prefix-free codes. Failing to do this leads to the rejected alternative UTF-Infinity.\n\n## Encoding\n\nThis section is based off the reference implementation which is written in Python. It is well documented, and is more specific on how to use bitwise operations. This is an abridged HTML version.\n\nSuppose that we have an unsigned integer `n` that we want to encode in UTF-8000. Initialize an empty dynamic array of bytes `ret_ints` that will store the UTF-8000 code unit.\n\nIf `n < 0x80`, eg `n = 0x41`, then insert `n` at the head of `ret_ints` and we are done. This is the ASCII byte for `n`, which in our example of `n = 0x41` is a capital letter a, `'A'`.\n\nOtherwise for `n ≥ 0x80`, eg `n = 0x0321C0FFEE8086`, we use UTF-8000. Initialize an integer counter `n_bits_content_occupied` to zero.\n\n                Our example `n`'s content bits look like `11 001000 011100 000011 111111 111011 101000 000010 000110` as a big raw number, with spaces added for visual ease.\n              \n\n                While `n` has more than 6 content bits, aka `n > 63 =` `00111111`, extract the least-significant 6 bits of `n` and insert them at the head of `ret_ints`, incrementing `n_bits_content_occupied` by 6 and downwards bitshifting `n` by 6.\n              \n\n                `n_bits_content_occupied = 48`, `n = 0b``11`,\n              \n\n`ret_ints`:\n                `00001000`\n                `00011100`\n                `00000011`\n                `00111111`\n                `00111011`\n                `00101000`\n                `00000010`\n                `00000110`\n              \n              Now insert the rest of `n` at the head of `ret_ints`. Count the number of bits left in `n` by downwards bitshifting `n` one bit at a time while it is non-zero. This is between 1 and 6 (inclusive), which we also add to `n_bits_content_occupied`.\n\n`n_bits_content_occupied = 50`, `n = 0`,\n\n`ret_ints`:\n                `00000011`\n                `00001000`\n                `00011100`\n                `00000011`\n                `00111111`\n                `00111011`\n                `00101000`\n                `00000010`\n                `00000110`\n              \n              Now we calculate how many bytes our UTF-8000 code unit requires, `n_utf_8000_bytes_needed`. We know that a `k`-byte code unit has capacity for `5k+1` content bits. Therefore `⌈(n_bits_content_occupied-1) / 5⌉` is the sufficient and minimal answer. Any larger code unit size would lead to an overlong encoding! For our example `n_utf_8000_bytes_needed = ⌈(50-1) / 5⌉ = 10`.\n\nLeftwards pad `ret_ints` with empty bytes to the length `n_utf_8000_bytes_needed`.\n\n`ret_ints`:\n                `00000000`\n                `00000011`\n                `00001000`\n                `00011100`\n                `00000011`\n                `00111111`\n                `00111011`\n                `00101000`\n                `00000010`\n                `00000110`\n              \n              \n                Now we add the start bits. The number of `1` start bits is equal to two less than the number of bytes in the code unit, which we just calculated. We therefore calculate `q, r = divmod(n_utf_8000_bytes_needed-2, 6)`, which tells us we need `q` hextets full of `1` start bits, and a final hextet of zero to five `1` bits, which also has space to contain the terminating `0` bit. In our example `(q = 1, r = 2) = divmod(10-2, 6)`.\n              \n\nApply the start bits across `ret_ints` using bitwise-or. The final start bits hextet can be given by `((1 << r) - 1) << (6 - r)`.\n\n`ret_ints`:\n                `00111111`\n                `00110011`\n                `00001000`\n                `00011100`\n                `00000011`\n                `00111111`\n                `00111011`\n                `00101000`\n                `00000010`\n                `00000110`\n              \n              Any of the lowest six bits of each byte that are not set by this point, unoccupied by content bits and untouched by start bits, are really content bits that the `k`-byte capacity provides but that we didn't need. Our example's `n_bits_content_occupied = 50` is one less than `5k+1 = 5*10+1 = 51`. We can color highlight it green as a content bit for completion's sake.\n\n`ret_ints`:\n                `00111111`\n                `00110011`\n                `00001000`\n                `00011100`\n                `00000011`\n                `00111111`\n                `00111011`\n                `00101000`\n                `00000010`\n                `00000110`\n              \n              \n                Now we crown the bytes with their self-synchronization prefixes, which delivers us from hextets to UTF-8000 octets. The first byte's self-synchronization prefix is\n                `11`, and continuation bytes have `10`.\n              \n\n`ret_ints`:\n                `11111111`\n                `10110011`\n                `10001000`\n                `10011100`\n                `10000011`\n                `10111111`\n                `10111011`\n                `10101000`\n                `10000010`\n                `10000110`\n              \n              And we're done!\n\n## Decoding\n\nAs with the encoding section, this section is based off the reference implementation which is written in Python and well documented.\n\nSuppose that we are receiving a stream of UTF-8000 bytes (possibly with errors!), and we wish to extract and taxonomically annotate the incoming code units. There are a few ways that we could approach this, such as using the byte map as a state machine, which I want to try in the future, or the classic way of using bitwise masks. We are going to use the latter approach in this section, as we describe how to decode a single code unit. But first, a look at error handling.\n\n### Error Recovery\n\nThe errors that can occur when decoding a UTF-8000 stream are:\n\n1. \n                  Reading a continuation byte (`10` ) when we are expecting the first byte of a code unit (`0` or`11` ).\n2. \n                  Reading a first byte (`0` or`11` ) when we are expecting a continuation byte (`10` ).\n3. Early EOF midway through a code unit.\n4. \n                  Encountering an overlong encoding\n                  \n  1. \n                      For 2-byte code units this is bytes 0xC0 (`11000000` ) and 0xC1 (`11000001` ).\n  2. \n                      For `n` -byte code units in general, with`n > 2` , eg`11100000``10010111``10010000` .\n5. \n                      For 2-byte code units this is bytes 0xC0 (\n6. Encountering an encoded surrogate codepoint in the range `U+D800` to`U+DFFF` , which is forbidden for compatibility with UTF-16.\n\nFor standard UTF-8 one would also have to be concerned with codepoints beyond the range `U+10FFFF` whence bytes 0xF5 to 0xFF go unused.\n\nFor any of these errors a parser could raise an exception and refuse to continue. Alternatively it could take advantage of UTF-8000's self-synchronization property, and keep calm and carry on\n\n, yielding Unicode replacement characters `U+FFFD` �\n\n until we reach the first byte of the next code unit. Let us investigate the latter course.\n\nTo handle error 1. the parser should return one � and get ready to parse the next code unit. When handling error 2. the parser should make sure to unpop\n\n the byte encountered, as it is the first byte of the next code unit. When handling errors 2. through to 5. there are a couple of mainstream approaches for yielding � characters:\n\n#### Maximal Subpart\n\n                The Unicode Consortium recommends, but does not enforce, a maximal subpart\n\n approach, in which the longest well-formed part of a code unit should return a single � character, rather than one for each byte involved. For example the three bytes in error 4.2. above should return one � as it is well-formed with respect to self-synchronization and self-punctuation, and only invalid at an overlong level, being an overlong encoding of\n                `11010111`\n                `10010000`\n                `U+05D0`, a Hebrew letter Aleph 'א'.\n              \n\n                I dislike this approach. Waiting for maximal subparts has the problem that the rest of an invalid code unit may never arrive. If we receive the bytes\n                `11100000`\n                `10010111`\n                from a socket, then the remote end may be waiting for us to chastise their overlong opening bytes with a response, because we can already tell that these bytes form part of an invalid code unit. Using the maximal subpart approach we *also* would be waiting, for the remote end to send a continuation byte eg `10010000` to form an overlong but otherwise complete 3-byte code unit. This is uncooperative, and not what I want.\n              \n\n#### One � For Each Byte Read\n\nWe are going to do what Python, my terminal KDE Konsole, and others do, and simply return a � character for each invalid byte. In Python `b'\\xE0\\x97\\x90'.decode(errors='replace')` returns `'���'`.\n\nThis approach is easier and more versatile. The end user can see how many invalid bytes occurred by counting the number of � characters. There are no deadlock waiting events that can occur with the maximal subpart approach.\n\n### The Main Decode Loop\n\nInitialize an empty dynamic array of bytes `parsed_bytes` that will store the bytes as we parse them.\n\nRead a byte, store it as `start_byte`. Use bitwise masks to find the index, `idx_0`, of the most-significant zero bit in the byte. If there are no zeros in this byte, `idx_0` should be set to `-1`.\n\n                If `idx_0 == 7` (`0xxxxxxx`) then `start_byte` is an ASCII byte, which has seven content bits. Append `start_byte` to `parsed_bytes` and we are done.\n              \n\n                If `idx_0 == 6` (`10xxxxxx`) then `start_byte` is a continuation byte, which is an invalid start byte. Go to error 1.\n              \n\n                If `idx_0 == 5` (`110xxxxx`) then this is the first byte of a 2-byte code unit. We treat this as a special case because there are only 4 mandatory content bits, not 5. As they are all contained in `start_byte` we can check them immediately for overlong encoding, to see if we need to handle error 4.1. If `start_byte` passes this check then append it to `parsed_bytes` and await a continuation byte. Handle error 2 if necessary, else append the continuation byte to `parsed_bytes` and we're done.\n              \n\nWe could (should really) make `idx_0 == 4` a special case too, to check for and forbid the surrogate ranges. I have omitted this for the time being and we drop through to the generic case below.\n\n                Otherwise we enter the generic case (`111[1...]`). Initialize an integer counter `n_bytes_expected` to 2. Increment `n_bytes_expected` by `5 - idx_0`, as `idx_0` now serves the purpose being the index of the terminating zero of the start bits, `0`.\n              \n\nIf `idx_0 == -1` then our code unit has multiple start bytes, exciting! Append `start_byte` to `parsed_bytes`, and `while(1)`:\n\n                Read a byte, and make sure it is a continuation byte lest we go to error 2. Use bitwise masks to find `idx_0`, the index of the most-significant zero bit in the lowest *six* bits of the byte, setting `idx_0` to `-1` if there is none. This is to continue trying to find the `0` start bit. Increment `n_bytes_expected` by `5 - idx_0`. If `idx_0 == -1` then append `start_byte` to `parsed_bytes` and continue again through this loop, until we find the `0` bit, at which point we break this loop.\n              \n\nAt this stage, whether our code unit has multiple start bytes or just one, `start_byte` is the *final* start byte of the code unit, `idx_0` is between 0 and 5 (inclusive), and we move towards checking for overlong encoding. Just as ordinals count the number of things less than themselves, `idx_0` counts the number of content bits contained `start_byte`, occupying the least significant bits.\n\nThere are six cases for anti-overlong checking, which correspond to `idx_0`'s value. That may sound like a lot, but the looping gif below that I made should relax you. It demonstrates periodic behavior. Even though it shows deep\n\n code unit sections with multiple start bytes, this animation still applies for *all* code units of length `n > 2`. The colored bars are based off the anatomy section image.\n\n                If `idx_0 == 5` then all the mandatory content bits are contained together in the final start byte. Thus we should immediately check `start_byte` using the mask `00011111`. We then read the first non-start byte, a continuation byte which does not need overlong checking (`10xxxxxx`).\n              \n\n                Otherwise we read another continuation byte, the first non-start byte. If `idx_0 == 0` then all the mandatory content bits are contained together in this first non-start byte (`10xxxxxx`), and we use the mask `00111110` to check for overlong encoding. Else `idx_0` is between 1 and 4 (inclusive) and the mandatory content bits are straddled across the final start byte and first non-start byte. In these cases we use two masks to check for overlong encoding, which one can see in the gif above.\n              \n\nPerhaps the case of `idx_0 == 0` could be grouped in with `idx_0` being between 1 and 4, by using an empty mask to check the final start byte, in order to make the algorithm less branch-y, but this walkthrough isolates which bytes are responsible for potential overlong encoding.\n\n                Given that the final start byte and first non-start byte have passed the overlong check, append them to `parsed_bytes`. Finally while the length of `parsed_bytes` is less than `n_bytes_expected`, read plain-old continuation bytes (`10xxxxxx`) and append them to `parsed_bytes`.\n              \n\nAnd we're done!\n\n## Further Ideas\n\n## Signed Variant: ZigZag Encoding\n\nSo far we have used UTF-8000 to encode codepoints, aka non-negative integers, aka unsigned integers. I have come up with a couple of modified interpretations of the content bits which allow us to encode the *entire* integers, aka the signed integers.\n\nWe make use of a marvelous bijective mapping between the signed integers and unsigned integers called the zigzag function\n\n that remains a bijection when restricting to the respective `n`-bit ranges. We use this as a final layer at the beginning/end of the standard UTF-8000 encoding/decoding procedure.\n\n### Source\n\n                  ZigZag encoding from Protobuf by Google: *Protocol Buffers Documentation / Encoding*\n                \n\nMyself: This seems like the perfect extensible solution for how to encode signed integers on top of UTF-8000.\n\n### TLDR / Examples\n\nThe code unit structure is identical to UTF-8000. The content bits correspond to the image of the zigzag function.\n\n| `zigzag(z)` | `z` | `UTF-8000` |  | \n|---|---|---|---|\n| ... | ... | ... |  | \n| `124` |  `62`   | `01111100` |  | \n| `125` | `-63` | `01111101` |  | \n| `126` |  `63` | `01111110` |  | \n| `127` | `-64` | `01111111` |  | \n| `128` |  `64` | `11000010` | `10000000` | \n| `129` | `-65` | `11000010` | `10000001` | \n| `130` |  `65` | `11000010` | `10000010` | \n| `131` | `-66` | `11000010` | `10000011` | \n| ... |  |  |  | \n\n### The ZigZag Function\n\nThe `zigzag` function maps from the signed integers to the unsigned integers.\n\nIf `z ≥ 0` then `zigzag(z) = 2 * z = (z << 1)`\n\nIf `z < 0` then `zigzag(z) = -2 * z - 1 = -(z << 1) - 1 = ~(z << 1)`\n\n| `z` | `zigzag(z)` | \n|---|---|\n| ... | ... | \n| `-4` | `7` | \n| `-3` | `5` | \n| `-2` | `3` | \n| `-1` | `1` | \n|  `0` | `0` | \n|  `1` | `2` | \n|  `2` | `4` | \n|  `3` | `6` | \n| ... | ... | \n\n| `zigzag(z)` | `z` | \n|---|---|\n| ... | ... | \n| `0` |  `0` | \n| `1` | `-1` | \n| `2` |  `1` | \n| `3` | `-2` | \n| `4` |  `2` | \n| `5` | `-3` | \n| `6` |  `3` | \n| `7` | `-4` | \n| ... | ... | \n\n```\n       ______________\n      /  __________  \\\n     /  /  ______  \\  \\\n    /  /  /  __  \\  \\  \\\n   /  /  /  /  \\  \\  \\  \\\n  -4 -3 -2 -1  0  1  2  3\n   \\  \\  \\  \\_____/  /  /\n    .  \\  \\_________/  /\n     .  \\_____________/\n      .\n                    \n```\n                  The ASCII art above illustrates the enumeration of the preimage of `zigzag`, showing it zigzagging between positives and negatives. This should make it clear how after `2^N` steps we have covered exactly the range `[-2^(N-1), 2^(N-1))`.\n\nFor example the preimage of the 7-bit unsigned range `[0, 128)` is the 7-bit signed range `[-64, 64)`.\n\n#### Branchless ZigZag\n\nIf one is dealing with fixed-width integers, for example mapping from `int32_t` to `uint32_t`, one can create a branchless version of `zigzag`, wow! CPUs like branchless code.\n\n                  With this example `zigzag(z) = (z << 1) ^ (z >> 31)`, where \n\n here is the C bitwise-xor operator.\n                `^`\n\nIf `2^31 > z ≥ 0` then `(z >> 31) = 0`, because we have filled the register with the highest bit of a non-negative signed number, 0. Thus `zigzag(z) = (z << 1) ^ (z >> 31)`. Okay, nothing special?\n\nBut if `-2^31 ≤ z < 0` then `(z >> 31) = -1`, because we have filled the register with the highest bit of a negative signed number, 1. Aha, so to achieve the bitwise complement, `~(z << 1)`, we can bitwise-xor with this `-1`. Thus `zigzag(z) = (z << 1) ^ (z >> 31)`.\n\n### Properties\n\nSelf-synchronization, self-punctuation, and arbitrary code unit length have the same conclusion as base UTF-8000. Properties that differ are discussed below.\n\n#### Encoded Range\n\nThe content bit counts work the same as they do for UTF-8000. Below is a summary of the ranges of integers that the content bits encode.\n\n| code unit length | number of content bits | minimum integer | maximum integer | \n|---|---|---|---|\n| `n = 1` | `7` | `-2 ^ (7n-1) ( = -64)` | `+2 ^ (7n-1) - 1 ( = +63)` | \n| `n ≥ 2` | `5n+1` | `-2 ^ (5n)`  | `+2 ^ (5n) - 1` | \n\n#### Small Integers, Small Code Units\n\nUTF-8000 is really just a variable-width bit container with some nice properties. Provided that we obey the forbidding of overlong encoding, we can use the content bits as we please, encoding from an arbitrary alphabet to unsigned integer codewords that form the content bits.\n\nThe alphabet in question for us is the signed integers, ℤ. We heuristically think of magnitude as a measure of commonness. The closer an integer is to zero, the more common it is, and thus the smaller the unsigned integer that it should be encoded as, whence the shorter the UTF-8000 code unit it occupies. This is almost common sense.\n\nWe have observed that `zigzag` achieves this. Two's complements in a fixed-width register however does *not* do this, as for example in a 64-bit CPU register the number -1 is encoded as `111...[64]...11`. This is not a problem for hardware like CPUs, but we are interested in efficient encoding for transmission and storage.\n\n#### Modified `strcmp(3)` Ordering\n\n                \n                  Since the negative integers are interwoven (zigzagged) between the non-negative integers via `zigzag`, we lose `strcmp` ordering from UTF-8000. For example -1 < 0 but `00000001` > `00000000`. However, being undeterred we can supersede this fact.\n                \n\nThe purpose of `strcmp(s1, s2)` with respect to UTF-8000 is to quickly compare code units `s1` and `s2` as a proxy for comparing their contained codepoints, without having to actually decode the code units. For this modified version of UTF-8000 we wish to create a quick proxy for comparing the contained signed integers.\n\nThe only variants of UTF-8000 that can make exact use of `strcmp` are those whose content bits encode letters from a totally-ordered alphabet, for which there exists an order-preserving bijection between that alphabet and the unsigned integers. Since the unsigned integers has a minimum element, 0, and the signed integers (our alphabet) does not have a minimum element, no such bijection exists.\n\nWe know that `zigzag(z)` breaks nicely into two cases, non-negative signed integers and negative signed integers. We also know that order-preserving bijections *do* exist between 0) non-negative signed integers and the even unsigned integers, and 1) negative signed integers and the odd unsigned integers. We initially break our new `strcmpsigned(s1, s2)` function into four cases depending on the final bit of each code unit, which we know is a content bit and indicates whether the stored unsigned integer is even or odd. We can obtain these bits via `b1 = c1 & 1` and `b2 = c2 & 1` where `c1` and `c2` are the final bytes of the respective code units:\n\n| `b1` | `b2` | `comment` | `b2 - b1` | `1 - b1 - b2` | \n|---|---|---|---|---|\n| `0` | `0` | `s1 ? s2` |  `0` |  `1` | \n| `0` | `1` | `s1 > s2` |  `1` |  `0` | \n| `1` | `0` | `s1 < s2` | `-1` |  `0` | \n| `1` | `1` | `s1 ? s2` |  `0` | `-1` | \n\nIf `b2 - b1` is non-zero then `strcmpsigned` can return that, as we are effectively comparing two signed integers of a different sign. Otherwise `(1 - b1 - b2) * strcmp(s1, s2)` effectively compares two integers of the same sign. We could therefore write this as:\n\n                  `strcmpsigned(s1, s2) = (b2 - b1) ? (b2 - b1) : (1 - b1 - b2) * strcmp(s1, s2)`\n                \n\nThat's pretty succinct! I have also assumed that `strcmp` is just the simple `{-1, 0, +1}` version.\n\n### Verdict\n\n                  I like it! I'll add it to the reference implementation when I get chance. It will most likely be a flag \n\n for `-z`\nzigzag\n\n used like `$ utf-8000 info -z -- -67` showing `11000010` `10000101`.\n                \n\nThis zigzag variant has a big advantage over the rejected two's complement signed variant in that we don't need to change how we do overlong checking from the standard UTF-8000 method. This means that we can effectively separate out into layers\n\n: 1) the overlong checking of code units and the extracting of their content bits, from 2) the further decoding of the content bits eg to a signed integer.\n\nThe only external metadata needed when decoding a stream of UTF-8000 bytes is is this unsigned or signed?\n\n. This is no different to decoding a stream of raw bytes, or inspecting fixed-width integers stored in two's complement form in a CPU register: it's up to the programmer's use-case to know whether signed or unsigned is expected.\n\n## UTF-16K\n\nCould we apply some techniques from this document to also extend UTF-16? Yes, but at the cost of forbidding more codepoints from being encoded similar to the forbidden surrogate range `U+D800` to `U+DFFF`.\n\nUTF-16 is a bit messy in its existing two-byte and four-byte form, but we can clean this up in a mostly forwards-compatible manner by using ASCVI-on-UTF-16. We assume the use of big-endian UTF-16 in this section.\n\n                  The high (`110110`) and low (`110111`) surrogate prefixes are highlighted in bright pink.\n                \n\nFor normal 20-content-bit surrogate pair UTF-16, the upper four content bits of high surrogates encode which Unicode Plane (collection of `2^16 = 64k` codepoints) that the code unit's content bits belong to. These bits are highlighted in bright crimson.\n\nFor multi-surrogate-pair UTF-16K we highlight only the upper three of these plane bits, the erstwhile fourth being a self-synchronization bit.\n\n### Source\n\nMyself: Realizing that I can challenge the UTF-16 extensions proposed by UCS-X.\n\n### TLDR / Examples\n\n| 2 | `xxxxxxxx xxxxxxxx` |  |  |  |  |  |  | \n| 4 | `110110xx xxxxxxxx` | `110111xx xxxxxxxx` |  |  |  |  |  | \n| 8 | `11011010 000xxxxx` | `110111xx xxxxxxxx` | `11011010 01xxxxxx` | `110111xx xxxxxxxx` |  |  |  | \n| 12 | `11011010 0010xxxx` | `110111xx xxxxxxxx` | `11011010 01xxxxxx` | `110111xx xxxxxxxx` | `11011010 01xxxxxx` | `110111xx xxxxxxxx` |  | \n| 16 | `11011010 00110xxx` | `110111xx xxxxxxxx` | `11011010 01xxxxxx` | `110111xx xxxxxxxx` | `11011010 01xxxxxx` | `110111xx xxxxxxxx` | ... | \n| 20 | `11011010 001110xx` | `110111xx xxxxxxxx` | `11011010 01xxxxxx` | `110111xx xxxxxxxx` | `11011010 01xxxxxx` | `110111xx xxxxxxxx` | ... | \n| ... |  |  |  |  |  |  |  | \n| 76 | `11011010 00111111` | `11011111 11111111` | `11011010 0110xxxx` | `110111xx xxxxxxxx` | `11011010 01xxxxxx` | `110111xx xxxxxxxx` | ... | \n| ... |  |  |  |  |  |  |  | \n| 144 | `11011010 00111111` | `11011111 11111111` | `11011010 01111111` | `11011111 11111111` | `11011010 01110xxx` | `110111xx xxxxxxxx` | ... | \n| ... |  |  |  |  |  |  |  | \n\n### Properties\n\n                  Two-byte UTF-16 is just the raw binary form of any 16-bit codepoint, except for the surrogate\n\n range `U+D800` to `U+DFFF` of size 2048 which any Unicode encoding (UTF-8, UTF-16, UTF-32) is forbidden to encode. The reason for this exclusion is because otherwise we could not distinguish `110110xx xxxxxxxx` from `110110xx xxxxxxxx` and `110111xx xxxxxxxx` from `110111xx xxxxxxxx` which you'll read about below.\n                \n\n                  Four-byte UTF-16 is two surrogate codepoints stuck next to each other, one high\n\n in the range `U+D800` to `U+DBFF` and then one low\n\n in the range `U+DC00` to `U+DFFF`. The real codepoint that they encode is 0x10000 added to the 20 binary digit number contained in the content bits within. For example `11011000 00001000` `11011100 00101101` contains 0x0202D, which then has 0x10000 added to it, to give 0x1202D. `U+1202D` is 𒀭\n\n, a Mesopotamian Dingir.\n                \n\n                  To expand UTF-16 indefinitely instead of being stuck with 0x110000 (1,114,112) codepoints, we employ ASCVI inside of UTF-16 surrogate pairs. UTF-8 was able to expand from 7-bit ASCII without any trouble because bytes with the highest bit set were undefined. In UTF-16 however every possible value that the content bits can take defines a codepoint. A smart choice, so that we do not interfere with already assigned codepoints, and so that we do not malapportion too many pre-existing unassigned codepoints for UTF-16K, and for there to be content bit count parity with UTF-8000, is to constrain ourselves to two unassigned planes, e.g. Planes 9 and 10. These planes are nice as the surrogate pairs take the form `11011010 0xxxxxx` `110111xx xxxxxxxx`. The codepoints `U+90000` to `U+AFFFF` are to be forbidden from being encoded, just as the 2048 surrogates of the Basic Multilingual Plane are.\n                \n\nWe use these four-byte surrogate pair containers as the quantum for UTF-16K, which uses two or more of these quanta to extend from UTF-16. The content bits encode the codepoint's binary representation, *without* adding on 0x10000 for the sake of simplicity, similar to UTF-8000.\n\n#### Bit Counts\n\n| number of bytes | number of content bits | number of mandatory content bits | \n|---|---|---|\n| `n = 2` | `16` | `0` | \n| `n = 4` | `20` | `0` | \n| `n = 8  ; k = 2` | `15k + 1 ( = 31)` | `10, or codepoint ≥ 0x110000` | \n| `n = 4k ; k ≥ 3` | `15k + 1` | `15` | \n\nOverlong checking is complicated for the jump from 4 bytes to 8 bytes, because of the 0x10000 that is added to the content bits. Either one of the highest 10 bits has a non-zero bit, or the `2^20` bit is active and one of the `2^m` bits is active with `16 ≤ m ≤ 19`.\n\nOne may notice when enumerating `15k + 1`, the number of content bits for `k`-surrogate-pair UTF-16K, that these values overlap predictably with UTF-8000's number of content bits given by `5n + 1`. This is the result of a deliberate choice to use two planes for UTF-16K instead of e.g. one plane, or half a plane etc.\n\n| number of UTF-16K surrogate pairs | number of UTF-8K bytes | number of content bits | \n|---|---|---|\n| `2` | `6` | `31` | \n| `3` | `9` | `46` | \n| `4` | `12` | `61` | \n| `5` | `15` | `76` | \n| ... | ... | ... | \n| `17` | `51` | `256` | \n| ... | ... | ... | \n\nThis means that UTF-8K and UTF-16K can be expanded in parallel in a predictable way, such that all possible codepoints from an expansion are permitted. This is in contrast to how 4-byte UTF-8 does not allow use of all `2^21` codepoints, but rather artificially restricts to `2^16 + 2^20` for parity with UTF-16K.\n\nThe number of bytes in such UTF-16K code units is always `4/3` that of an equivalent UTF-8K code unit. The reciprocal of this, `3/4`, ends up as the scaling factor of the information rate limit from UTF-8K to UTF-16K. It is nice that this ratio is independent of code unit length.\n\n#### Information Rate\n\nFor 2-byte UTF-16 this is technically `16 / 16 = 100%`, ignoring the forbidden surrogate range.\n\nFor 4-byte UTF-16 this is `20 / 32 = 62.5%`, again ignoring the forbidden UTF-16K planes 9 and 10. This looks slightly worse than UTF-8's `21 / 32`, but do bear in mind that UTF-8 is also restrained to UTF-16's upper limit.\n\nFor beyond four bytes this is `(15k+1) / (4k*8) = 15/32 + 1/(32k)` which approaches `15/32 = 46.875%`, which is okay. As predicted above, this is `3/4` times the information rate limit of UTF-8000. `3/4 * 5/8 = 15/32`.\n\nBelow is a comparison of the efficiencies of UTF-8 and UTF-16.\n\n| start | end | range size | number of UTF-8 bytes | number of UTF-16 bytes | \n|---|---|---|---|---|\n| `U+0000` | `U+007F` | `0x80 = 128` | `1 (ASCII)` | `2` | \n| `U+0080` | `U+07FF` | `0x780 = 1,920` | `2` | `2` | \n| `U+0800` | `U+FFFF` | `0xF800 = 63,488` | `3` | `2` | \n| `U+10000` | `U+10FFFF` | `0x100000 = 1,048,576` | `4` | `4` | \n\nUTF-16 is more efficient than UTF-8 only at encoding `U+0800` to `U+FFFF`, aka the three-byte UTF-8 range that UTF-16 encodes using two bytes.\n\nFor ASCII, and beyond `U+10FFFF`, UTF-8000 is far more efficient (and less ugly) than UTF-16K.\n\n#### Self-Synchronization\n\nUTF-16K does not have self-synchronization at the byte level because UTF-16 does not. If one experiences a single missing byte then potentially the whole stream becomes corrupted.\n\n                  At the two-byte level UTF-16 has self-synchronization which UTF-16K inherits. Non-surrogate codepoints are quantum; they are to UTF-16 as ASCII is to UTF-8. Surrogate pairs provide self-synchronization with their `110110` and `110111` high and low surrogate prefixes.\n                \n\n                  UTF-16K goes even deeper, using multiple surrogate pairs that shadow Planes 9 and 10. Within the surrogate pair self-synchronization level, within the high surrogates used to encode UTF-16K, `11011010 0xxxxxx`, ASCVI is employed, whose leading bit provides self-synchronization, with `11011010 00` and `11011010 01`.\n                \n\nTherefore overall UTF-16K exhibits self-synchronization at the two-byte level, like UTF-16.\n\n#### Self-Punctuation\n\nInherited from ASCVI.\n\n#### Compatibility\n\nUTF-16K forbids Planes 9 and 10 of Unicode, because it has no way to encode those codepoints, instead repurposing the surrogate pairs erstwhile required to encode Planes 9 and 10 for the purpose of encoding codepoints beyond 0x110000. An important question to ask regarding forwards compatibility is what does an existing UTF-16 parser do if it meets a UTF-16K code unit?\n\n.\n\nIn short it's Plane-9-or-10-garbage-in Plane-9-or-10-garbage-out. Each surrogate pair used in encoding a UTF-16K codepoint beyond 0x110000 would be parsed separately as though it belongs to Plane 9 or 10, but with no *syntactic* issues. *Semantically* however this would cause issues with logical character (codepoint) counts that would count each surrogate pair as a separate character, rather than contributing towards a single character.\n\nI think that this is a better solution than UCS-X's UTF-G-16 which breaks syntactic compatibility with UTF-16 by repurposing low surrogates as leading units\n\n for UTF-G-16 6-byte code units. One could argue that UTF-G-16 is better because those bytes could be replaced with a single replacement character �\n\n though I'm not convinced, as for example the default behavior of Python's `bytes.decode` function is `'strict'`, which raises an exception, not `'replace'` which produces replacement characters. UTF-G-16 also has flawed error handling behavior as discussed in the feedback emails, arising from UTF-G-16's self-synchronization requiring a context-dependent interpretation of low surrogates to determine if they are leading\n\n or trailing\n\n, whereas UTF-16K's self-synchronization is context-independent by using a disjoint union of planes 9 and 10.\n\nThe requirement to extend the list of codepoints that all of UTF-8, UTF-16, and UTF-32 are forbidden from encoding, to include Planes 9 and 10 or elsewhere, would not be an easily negotiated feat. We would be banning an extra `2/17 = 11.8%` of pre-existing codepoints. One may notice that this situation of forbidding pre-existing codepoints is a similar situation to back when 2-byte UTF-16 extended to 4-byte UTF-16. Would we ever have to ban codepoints in pre-existing ranges again after this UTF-16K extension? No, as UTF-8K and UTF-16K are infinitely extensible.\n\n### Verdict\n\n                  The immature part of me says let UTF-16 decay and die as the short-sighted, legacy, Windows, \n\n. But it will be around for a while, with several uses, like the Joliet Filesystem for my beloved Arch Linux ISOs grrr.\n                `wchar_t`, non-self-synchronizing-at-the-byte-level, +0x10000, garbage that it is\n\nThe main reason I wrote this section was to provide an alternative to UCS-X, so that we can use the same(ish) style as UTF-8000, and lest UCS-X or an even uglier idea come along.\n\nI predict that the Unicode Consortium would heavily push back on the idea of having to ban more codepoints, `U+90000` to `U+AFFFF`. No matter how one plans to extend UTF-16, it requires either forbidding some codepoints, or changing the syntax, either of which is a breaking change.\n\nFor our modern times UTF-8 is undoubtedly the way to go, and by the time that we need to extend to UTF-8000, I would hope that UTF-16 and all other encodings belong in a museum, and we can therefore extend UTF-8 without worrying about compatibility with the others.\n\n### Reference Implementation\n\nAvailable! See below.\n\n## UTF-32K\n\nIn the same manner that ASCII is extended by UTF-8 and UTF-8000, UTF-32 could also be extended to be a multi-\nbyte\n\n (32-bit chunk) encoding scheme.\n\nDo we really want this though? Is UTF-32 meant to be variable-width, or is it meant to represent the raw codepoint, decoded and stored in memory as a fixed-width integer? I've written this section to demonstrate that ASCVI can be applied to a quantum as small as 3 bits (seriously lol), or large like 32 bits.\n\n### Source\n\nMyself: It seemed obvious how this follows from UTF-8000.\n\n### TLDR / Examples\n\nWe could extend UTF-32 either in the style of UTF-8000, treating the one-\nbyte\n\n code units as a special case occupying the lower 31 bits...\n\n| 1 | `0xxxxxxx_xxxxxxxx_xxxxxxxx_xxxxxxxx` |  | \n| 2 | `110xxxxx_xxxxxxxx_xxxxxxxx_xxxxxxxx` | `10xxxxxx_xxxxxxxx_xxxxxxxx_xxxxxxxx` | \n| ... |  |  | \n\n...or in the style of ASCVI, with no special cases and the one-\nbyte\n\n code units occupying the lower 30 bits.\n\n| 1 | `00xxxxxx_xxxxxxxx_xxxxxxxx_xxxxxxxx` |  | \n| 2 | `010xxxxx_xxxxxxxx_xxxxxxxx_xxxxxxxx` | `1xxxxxxx_xxxxxxxx_xxxxxxxx_xxxxxxxx` | \n| ... |  |  | \n\n### Properties\n\nMutatis mutandis, the properties of UTF-8000 and ASCVI apply. We only make further remarks on a couple of properties.\n\n#### Bit Counts\n\nPredictable like ASCVI.\n\n| number of bytes | number of content bits | number of mandatory content bits | \n|---|---|---|\n| `n = 4` | `30` | `0` | \n| `n = 4k ; k ≥ 2` | `30k` | `30` | \n\n#### Information Rate\n\nWhilst the ASCVI-style UTF-32K has an information rate of `30 / 32 = 93.75%`, it is very inefficient for low-value codepoints, with the highest bytes most often being zeros. UTF-8000 has a much finer telescopic expansion mechanism at the byte level, compared to UTF-32 at a four-byte level.\n\n#### Self-Synchronization\n\nUTF-32 is not self-synchronizing at the byte level, and UTF-32K inherits this weakness. UTF-32 is only self-synchronizing at the four-byte level, similar to how UTF-16 is only self-synchronizing at the two-byte level. UTF-32K maintains self-synchronization at the four-byte level.\n\n#### Endianness\n\nLike UTF-16, and unlike UTF-8 and UTF-8000, UTF-32 has endianness, its quantum being a whopping four bytes.\n\n### Verdict\n\nNot our greatest priority.\n\nUTF-32 is hardly ever used for transmission or storage due to its inefficiency and endianness.\n\nAs I wrote in the intro, UTF-32's main use is as a **non**-variable-width container, for when one decodes UTF-8 or UTF-16 to `int32_t` integers (UTF-32) for use inside a program. UTF-32K would be an anti-pattern / counterproductive.\n\n## Rejected Alternatives\n\nAlthough UTF-8000 extends naturally\n\n from UTF-8, is it still the best approach? Are there any better alternatives that engineer\n\n an extension from UTF-8, just as UTF-8 engineers an extension from ASCII?\n\nWe rule out a few alternatives in this section. It's good to document the suboptimal solutions (and outright failures) so that we can work towards success. I've done that plenty of times with my own ideas don't worry! Feel satisfied in having at least made an attempt.\n\n## ASCVI\n\nWhat if ASCII were only six-bit instead of seven-bit? Would this make extending to multi-byte code units more pleasant?\n\n### Source\n\nMyself: The intuitive derivation section of UTF-8000.\n\n### TLDR / Examples\n\n| 1 | `00xxxxxx` |  |  |  |  |  | \n| 2 | `010xxxxx` | `1xxxxxxx` |  |  |  |  | \n| 3 | `0110xxxx` | `1xxxxxxx` | `1xxxxxxx` |  |  |  | \n| ... |  |  |  |  |  |  | \n| 17 | `01111111` | `11111111` | `1110xxxx` | `1xxxxxxx` | ... | `1xxxxxxx` | \n| ... |  |  |  |  |  |  | \n\n### Properties\n\nSelf-synchronization, self-punctuation, `strcmp` ordering, BOM support, and arbitrary code unit length have the same conclusion as UTF-8000. Properties that differ are discussed below.\n\n#### Bit Counts\n\nThe number of content bits and mandatory content bits are even more predictable than those of UTF-8.\n\n| code unit length | number of content bits | number of mandatory content bits | \n|---|---|---|\n| `n = 1` | `6n ( = 6)` | `0` | \n| `n > 1` | `6n` | `6` | \n\n                  This is because one-byte code units are not special. They use the same `0` self-synchronization prefix as any first byte. The number 6 arises from every subsequent continuation byte adding on 7 more content bits, minus 1 for the longer start bits sequence.\n                \n\nConsequently the number of content bits stored in an `n` byte code unit is never a power of two, unlike with UTF-8000. This is because 6, containing 3 in its prime factorization, cannot divide into a power of two. This is not a terrible defect, but we do like powers of two.\n\n#### Information Rate\n\nASCVI's information rate is `6n / 8n = 6 / 8 = 75%`. This a constant independent of the length of the code unit.\n\nFor one-byte code units UTF-8000 (ASCII) is more efficient and versatile, storing double the number of codepoints and having an information rate of `87.5%`.\n\nFor multi-byte code units ASCVI is more efficient, with UTF-8000's information rate tending downwards towards `62.5%`.\n\nEven if the US English alphabet had its 52 letters cut down to eg 27 Hebrew glyphs, or no letters at all, one would struggle to create a practical set of 64 glyphs for single-byte ASCVI. The tradeoff of ASCII being seven-bit, at the slight detriment of the information rate of multi-byte code units, seems worth it.\n\n#### Byte Map\n\n|  | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | A | B | C | D | E | F | \n|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|\n| 0 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | \n| 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | \n| 2 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | \n| 3 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | \n| 4 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | \n| 5 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | \n| 6 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | \n| 7 | 4 | 4 | 4 | 4 | 4 | 4 | 4 | 4 | 5 | 5 | 5 | 5 | 6 | 6 | 7 | 8+ | \n| 8 |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  | \n| 9 |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  | \n| A |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  | \n| B |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  | \n| C |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  | \n| D |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  | \n| E |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  | \n| F |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  | \n\nLook at that beautiful geometric series layout. Four rows for `1`, two rows for `2`, one row for `3`, half a row for `4` etc...\n\nUnlike UTF-8000 which can never use the bytes 0xC0 and 0xC1, ASCVI uses all 256 possible bytes. Which of these is the advantageous behavior depends on whether one wants to do something extraordinary with those two bytes, or one wants to dissuade their use in chicanery.\n\n### Intuitive Derivation\n\nSee the intuitive derivation of UTF-8000 for how one might come up with this.\n\n### Naming\n\nThe II\n\n in ASCII reminds me of the Roman Numerals VII\n\n for seven, and ASCII is a seven-bit code. Therefore for this six-bit code we choose to use the Roman Numerals for six, VI\n\n, and name it ASCVI\n\n.\n\n### Verdict\n\n7 bit ASCII, and UTF-8 that extends it, are very well established. One-byte ASCVI is inferior to the flexibility of ASCII, albeit this contributes to UTF-8 having a slightly lower information rate for multi-byte code units. I do not yearn for an alternate universe, or a fresh start of text encoding standards, where ASCII is six-bit instead of seven.\n\nThat being said, ASCVI is by no means *inherently* flawed, and we can make use of it in UTF-16K and UTF-32K. In a sense ASCVI is the Platonic Form of UTF-8.\n\n## Signed Variant: Two's Complement\n\nOne might be shocked to find our beloved two's complement representation of the signed integers present in the rejected ideas section. This is not about one's personal taste in representing signed integers, but technological extensibility, with which the zigzag signed variant far outshines this two's complement variant.\n\n### Source\n\nMyself: It seemed like a good(ish) idea until I realized that the zigzag variant is better.\n\n### TLDR / Examples\n\nWe treat the `n` content bits of a code unit as a two's complement signed form, where the highest bit no longer has value `2 ^ (n-1)` but rather `-2 ^ (n-1)`.\n\n| ASCII |  |  |  |  |  | \n| 1 | `0xxxxxxx` |  |  |  |  | \n| UTF-8 |  |  |  |  |  | \n| 2 | `110xxxxx` | `10xxxxxx` |  |  |  | \n| 3 | `1110xxxx` | `10xxxxxx` | `10xxxxxx` |  |  | \n| 4 | `11110xxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  | \n| UTF-8000 |  |  |  |  |  | \n| 5 | `111110xx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | \n| ... |  |  |  |  |  | \n\nThe difference in code unit layout from UTF-8000 is that the mandatory content bits are downshifted by one bit.\n\n### Properties\n\n#### Overlong Encoding Checking\n\nPreventing overlong encoding requires checking that the content bits of an `n`-byte code unit encode an integer `z` from the range `[ -2^(5n), +2^(5n) ) \\ [ -2^(5(n-1)), +2^(5(n-1)) )`. This may look complicated, but we can break it down into two cases:\n\n                  For `z ≥ 0`, the highest content bit of both `n`-byte and `n-1`-byte code units is `0`. There must be at least one `1` bit in the following bits, which are the mandatory content bits.\n                \n\n                  For `z < 0`, the highest content bit of both `n`-byte and `n-1`-byte code units is `1`. This is to make `z` negative, using the `-2^(5n)` bit. There must be at least one `0` bit in the following bits, which are the mandatory content bits. Otherwise if all of these bits were ones, then `z` would be at least `-2^(5(n-1))`. For example the overlong encoding\n                  `11011111`\n                  `10000000` encodes `z = -64` which fits into a 1-byte code unit. Encoding integers below `-64` requires subtracting from these content bits, which sets at least one of the mandatory content bits to zero.\n                \n\n                  UTF-8000 never uses the bytes 0xC0 and 0xC1, which is explained in the glossary section for overlong encoding. Slightly differently, this two's complement signed variant never uses the bytes 0xC0 (`11000000`) or 0xDF (`11011111`).\n                \n\n#### No `strcmp(3)` Ordering\n\n                \n                  Since negative integers set the highest content bit to `1`, we lose `strcmp` ordering from UTF-8000. For example -1 < 0 but `01111111` > `00000000`.\n                \n\n### Verdict\n\n                  A big downside of this two's complement signed variant is that its anti-overlong checking mechanism differs from that of UTF-8000 because of the 1-bit-downshifted position of the mandatory content bits. Consequently it is not possible to agnostically decode a stream of these bytes as though they were UTF-8000 bytes. For example, a 0xC1 byte is valid in this variant as `11000001`, but is invalid in UTF-8000 as `11000001`.\n                \n\nWhat if we tried to redeem this variant by proposing to move the negative\n\n bit, the `-2^(5n)` bit, to the stable tail end of the code unit, rather than it being at the ever-expanding head of the code unit? This could hopefully mean that we would not have to change the anti-overlong checking mechanism from UTF-8000. The long and short is that we would perchance intuitively reinvent the zigzag signed variant from first principles, which indeed leftwards bitshifts by one the signed integer that it encodes, using the lowest bit as the negative\n\n bit. Our redemption is found there.\n\n## UTF-Infinity\n\nWhat if we naively roll the start bits\n\n over into further bytes?\n\n### Source\n\n                  Mashpoe on YouTube: *Expanding the UTF-8 Character Set to Infinity*\n                \n\n### TLDR / Examples\n\n| ... |  |  |  |  |  |  | \n| 7 | `11111110` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` |  | \n| 8 | `11111111` | `0xxxxxxx` | `10xxxxxx` | ... | `10xxxxxx` |  | \n| 9 | `11111111` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` |  | \n| 10 | `11111111` | `110xxxxx` | `10xxxxxx` | ... | `10xxxxxx` |  | \n| ... |  |  |  |  |  |  | \n| 15 | `11111111` | `11111110` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` | \n| 16 | `11111111` | `11111111` | `0xxxxxxx` | `10xxxxxx` | ... | `10xxxxxx` | \n| 17 | `11111111` | `11111111` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` | \n| 18 | `11111111` | `11111111` | `110xxxxx` | `10xxxxxx` | ... | `10xxxxxx` | \n| ... |  |  |  |  |  |  | \n\n### Properties\n\n#### Bit Counts\n\nEffectively, every jump from `8k-1`-byte code units to `8k`-byte code units the encoding inserts another blank byte after the start bytes, by which 8 minus 1 equals 7 bits of content are gained in an ASCII-looking byte, instead of appending a continuation byte by which 6 minus 1 equals 5 bits of content are gained.\n\n| code unit length | number of content bits | number of mandatory content bits | \n|---|---|---|\n| `n = 1` | `7` | `0` | \n| `n ≠ 8k` | `5n+1 + 2⌊n/8⌋` | `5` | \n| `n = 8k` | `5n+1 + 2⌊n/8⌋` | `7` | \n\n#### Information Rate\n\nFor an `n`-byte code unit the information rate is UTF-8000's information rate plus `2⌊n/8⌋ / (8n)`.\n\nI'm not working it out fully, but I can tell that this leads to a sawtooth-y profile as a graph of information rate against `n`. Therefore, counterintuitively, longer code units can have better efficiency than shorter ones.\n\n#### No Self-Synchronization\n\n                  In the 8-byte code unit example, there is no way to distinguish the second byte `0xxxxxxx` from an ASCII byte `0xxxxxxx`. This generalizes to `8n`-byte code units.\n                \n\n                  In the 15-byte code unit example, there is no way to distinguish the second byte, `11111110` from the first byte of a 7-byte code unit. This generalizes to `8n-1`-byte code units.\n                \n\n                  In the 16-byte code unit example, there is no way to distinguish the second byte, `11111111` from the first byte of an 8-byte code unit. This generalizes such that if one seeks to any `11111111` byte, one has no idea if this is the first byte of a code unit or not.\n                \n\nThis list is non-exhaustive.\n\n#### Self-Punctuation\n\nThis is the property that Mashpoe clearly prioritized preserving, however the approach was too myopic and did not lead to preserving other properties of interest.\n\n#### Patented\n\n\n                Mashpoe jokes (?) in the video that he owns the patent to this encoding scheme.\n\nRegardless of whether he is joking or not, I nonetheless find it reprehensible that someone *could* (at least try to) copyright / patent the correct way to extend UTF-8. It would be like trying to copyright the right solution to a mathematics equation, or a prime number! Therefore I am being quite loud in the copylefting of UTF-8000 in the licensing section. Everyone benefits from shared, free-as-in-freedom, open ideas.\n\n### Verdict\n\nThe loss of self-synchronization is a fatal detriment.\n\nThe formula for the number of content bits has predictable but irritable jumps, which lead to counterintuitive information rates.\n\n## Perl utf8\n\n                  Use *up to* 7 bytes to encode up to 36 bits of information in the sane way, in order to encode 32-bit integers (and a little beyond). To encode 64-bit integers, use a special-case *fixed* 13-byte code unit starting with `11111111`.\n                \n\nSince the `n`-byte code units with `n < 8` are the same as UTF-8000 we shall mostly only discuss the 13-byte code units.\n\n### Source\n\n                  Larry Wall for Perl5 on GitHub: *utf8.h*\n                \n\nA comment reads: A note on nomenclature: The term UTF-8 is used loosely and inconsistently in Perl documentation ... perl uses an extension of UTF-8 to represent code points that Unicode considers illegal.\n\n.\n\n### TLDR / Examples\n\n| ... |  |  |  |  |  |  |  |  |  |  | \n| 5 | `111110xx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  |  |  |  |  | \n| 6 | `1111110x` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  |  |  |  | \n| 7 | `11111110` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  |  |  | \n| 13 | `11111111` | `10000000` | `10000xxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` | \n\n### Properties\n\n#### Bit Counts\n\nThere are no code units of length 8, 9, 10, 11, or 12. Nor are there any of length 14 or beyond.\n\n| code unit length | number of content bits | \n|---|---|\n| `n = 1` | `7` | \n| `1 < n < 8` | `5n+1` | \n| `n = 13` | `63/64 used, up to 72?` | \n\n#### Why 13 Bytes Instead Of 12?\n\n                  Encoding the maximum possible signed 64-bit integer, 0x7FFF_FFFF_FFFF_FFFF, in Perl utf8 returns `11111111` `10000000` `10000111` ... `10111111`. Even if the maximum possible unsigned 64-bit integer, 0xFFFF_FFFF_FFFF_FFFF, were encodable it could fit into the lower 11 bytes. So why use 13 bytes instead of 12, with the second byte always `10000000`?\n                \n\n                  My guess is that the designers were going to use a UTF-Infinity style `11111111` `11111000` opening, but then they recognized that they would lose self-synchronization because the second byte looks like the start of a 5-byte utf8 sequence. Therefore they swapped the second byte to a `10000000`. They could have removed it and used 12 bytes, but that would preclude any future reunification with e.g. UTF-8000, which with 12 bytes has only `5 * 12 + 1 = 61` content bits which is less than 64, but with 13 bytes has `5 * 13 + 1 = 66` which is sufficient.\n                \n\n#### Information Rate\n\nOn the surface 13-byte, 72-bit code units have an information rate of `72 / (13*8) = 72 / 104 = 69%`, which is better than that of UTF-8000 (`62.5%`).\n\nHowever as only 64 of those 72 bits are used in encoding 64 bit numbers, with the whole of the first continuation byte never being used, the information rate is closer to `64 / (13*8) = 64 / 104 = 61.5%`, which is worse than UTF-8000.\n\n#### Self-Synchronization\n\n                  This is effectively the same as UTF-8000. All continuation bytes have a `10` self-synchronization prefix, and the 13-byte start byte `11111111` has a `11` self-synchronization prefix.\n                \n\n#### Self-Punctuation\n\n                  The first byte of a 13-byte code unit being `11111111` characterizes it as a special case, providing self-punctuation. This is similar to ASCII being a special case with its characterizing prefix of `0` in the highest bit.\n                \n\n#### Not Infinitely Extensible\n\nBecause Perl only supports up to 64-bit numbers without a specialized `bigint` module, it was sensible of them to cap their extension of UTF-8 to a finite number of bytes. It's not the prettiest however, and I'm not sure why they chose 13 bytes when 12 would suffice. CPU alignment if they don't store the predictable start byte of all 1s?\n\n### Verdict\n\nInextensible, providing only one special case beyond 7-byte UTF-8\n\n to encode 64-bit numbers, and is thus not widely known or supported.\n\n## UCS-X\n\nUCS-X proposes three extensions for each of UTF-8, UTF-16, UTF-32, for a total of nine specifications, twelve including the existing base specifications!\n\nI have so far only investigated the UTF-8 extensions, as they are all dense reads, and that is what we summarize in this section, with our main contribution being bitwise color highlighting.\n\nAt a glance the UTF-16 extensions look like they break syntax with base UTF-16, whereas our UTF-16K proposal does not. Ours only semantically reinterprets the high surrogates `U+DB00` to `U+DB3F`. I will have a look at UCS-X's UTF-16 and UTF-32 extensions when I get time, to see if they contain anything interesting, or if I'm wrong.\n\n### Source\n\n                  Tom Bishop and Richard Cook on ucsx.org: *The UCS-X Family of UCS Extensions (Draft Proposal)*\n                \n\n| UTF-8 | UTF-16 | UTF-32 | \n|---|---|---|\n| UTF-G-8 | UTF-G-16 | UTF-G-32 | \n| UTF-E-8 | UTF-E-16 | UTF-E-32 | \n| UTF-∞-8 | UTF-∞-16 | UTF-∞-32 | \n\n### TLDR / Examples\n\n#### UTF-G-8\n\nThe same as original 6-byte UTF-8 (RFC 2279) by Ken Thompson and Rob Pike, the same as UTF-8000.\n\n| ... |  |  |  |  |  |  | \n| 5 | `111110xx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  | \n| 6 | `1111110x` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | \n\n#### UTF-E-8\n\nThe same as Perl utf8.\n\n| ... |  |  |  |  |  |  |  |  |  |  | \n| 7 | `11111110` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` |  |  |  | \n| 13 | `11111111` | `10000000` | `10000xxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` | \n\n#### UTF-∞-8\n\nThis extension's code units are best characterized by the number of hex digits that the `U+...XXXX` codepoint representation consists of. The idea is that by adding on two more continuation bytes, which contain 12 content bits, one can add three more hex digits to the `U+...XXXX` codepoint representation.\n\nThe 18 hex digit, 71 and 72 content bit cases are handled specially, in the transition region of extending from UTF-E-8.\n\n                  Otherwise, to encode an integer N: Subtract 18 from the number of hex digits in the integer's `U+...XXX` codepoint representation. Store this number in one or more low length-storage bytes\n\n of the form `1010xxxx`. Precede these low length-storage bytes\n\n with the constant high length-storage bytes\n\n `10110100` (0xB4), where the number of high length-storage bytes is one less than the number of low length-storage bytes. Now precede this with the constant full start byte `11111111`. Now succeed all of this with continuation bytes that store the content bits, of the form `10xxxxxx`. These content bytes come in pairs, and the number of pairs should be one third of the number of hex digits, rounded up to the next integer if necessary.\n                \n\n| hex digits | content bits | bytes |  |  |  |  |  |  |  |  |  |  |  | \n|---|---|---|---|---|---|---|---|---|---|---|---|---|---|\n| ... |  |  |  |  |  |  |  |  |  |  |  |  |  | \n| 18 | 71 | 13 | `11111111` | `100xxxxx` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` |  |  |  |  |  | \n| 18 | 72 | 14 | `11111111` | `10100000` | `101xxxxx` | `10xxxxxx` | ... | `10xxxxxx` |  |  |  |  |  | \n| 19 | 76 | 16 | `11111111` | `10100001` | `10000000` | `1000xxxx` | `10xxxxxx` | ... | `10xxxxxx` |  |  |  |  | \n| 20 | 80 | 16 | `11111111` | `10100010` | `100000xx` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` |  |  |  |  | \n| 21 | 84 | 16 | `11111111` | `10100011` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` |  |  |  |  | \n| ... |  |  |  |  |  |  |  |  |  |  |  |  |  | \n| 31 | 124 | 24 | `11111111` | `10101101` | `10000000` | `1000xxxx` | `10xxxxxx` | ... | `10xxxxxx` |  |  |  |  | \n| 32 | 128 | 24 | `11111111` | `10101110` | `100000xx` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` |  |  |  |  | \n| 33 | 132 | 24 | `11111111` | `10101111` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` |  |  |  |  | \n| 34 | 136 | 28 | `11111111` | `10110100` | `10100001` | `10100000` | `10000000` | `1000xxxx` | `10xxxxxx` | ... | `10xxxxxx` |  |  | \n| 35 | 140 | 28 | `11111111` | `10110100` | `10100001` | `10100001` | `100000xx` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` |  |  | \n| 36 | 144 | 28 | `11111111` | `10110100` | `10100001` | `10100010` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` |  |  | \n| ... |  |  |  |  |  |  |  |  |  |  |  |  |  | \n| 271 | 1084 | 186 | `11111111` | `10110100` | `10101111` | `10101101` | `10000000` | `1000xxxx` | `10xxxxxx` | ... | `10xxxxxx` |  |  | \n| 272 | 1088 | 186 | `11111111` | `10110100` | `10101111` | `10101110` | `100000xx` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` |  |  | \n| 273 | 1092 | 186 | `11111111` | `10110100` | `10101111` | `10101111` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` |  |  | \n|  |  |  | `11111111` | `10110100` | `10110100` | `1010xxxx` | `1010xxxx` | `1010xxxx` | `10000000` | `1000xxxx` | `10xxxxxx` | ... | `10xxxxxx` | \n|  |  |  | `11111111` | `10110100` | `10110100` | `1010xxxx` | `1010xxxx` | `1010xxxx` | `100000xx` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` | \n|  |  |  | `11111111` | `10110100` | `10110100` | `1010xxxx` | `1010xxxx` | `1010xxxx` | `10xxxxxx` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` | \n| ... |  |  |  |  |  |  |  |  |  |  |  |  |  | \n\n                  The maximum integer that `n` low length-storage bytes can store is `16^n - 1`, which is always congruent to `0 mod3`. Thus the number of code units in each family of `n` low length-storage bytes is `(16^n - 1) - (16^(n-1) - 1)` which is also always congruent to `0 mod3`. In the base case of `n = 1`, the maximum low length-storage byte `10101111` (33 hex digits) is succeeded by the bytes\n                  `10xxxxxx`\n                  `10xxxxxx`, case 3/3 of the repeating pattern of mandatory content bits placement in the highest two content bytes. Thus we conclude inductively that each family of code units with `n` low length-storage bytes ends the same way, with\n                  `10101111`\n                  [ `10101111` ... ]\n                  `10xxxxxx`\n                  `10xxxxxx`. This ensures clean transitions from `n` to `n+1` low length-storage byte families.\n                \n\n### Properties\n\n#### Bit Counts\n\n| variant | code unit length | number of content bits | same as | \n|---|---|---|---|\n| ASCII | `n = 1` | `7` | UTF-8000 | \n| UTF-8 | `2 ≤ n ≤ 4` | `5n+1` | UTF-8000 | \n| UTF-G-8 | `5 ≤ n ≤ 6` | `5n+1` | UTF-8000 | \n| UTF-E-8 | `n = 7` | `5n+1 ( = 36)` | UTF-8000 | \n| UTF-E-8 | `n = 13` | `63` | Perl utf8 | \n| UTF-∞-8 | `n = 13` | `71` |  | \n| UTF-∞-8 | `n = 14` | `72` |  | \n\nand then further for UTF-∞-8:\n\n| number of hex digits | number of content bits | code unit length | \n|---|---|---|\n| `n ≥ 19` | `4n` | `2(⌊log`<sub>16</sub> (n-18)⌋ + 1) + 2⌈n/3⌉ | \n\nUTF-G-8 stores 31 bits, enough to encode positive signed 32-bit integers.\n\nUTF-E-8 stores 63 bits, enough to encode positive signed 64-bit integers.\n\nFor UTF-∞-8 this is unlimited.\n\n#### Information Rate\n\n                  Asymptotically `4n / (8 * 2(⌊log` tends towards <sub>16</sub>(n-18)⌋ + 1) + 2⌈n/3⌉)`3/4 = 75%`, which is better than that of UTF-8000 (`62.5%`). That's the power of the low length-storage bytes `1010xxxx` using all possible combinations of bits, whereas UTF-8000's start bits use only unary codewords. The `3/4 = 6/8` is representative of the content bytes.\n                \n\n#### Self-Synchronization\n\n                  Every byte beyond the first begins with the continuation prefix `10`, ensuring self-synchronization.\n                \n\n#### Self-Punctuation\n\n                  The high length-storage bytes `10110100` provide self-punctuation. They tell us to keep reading a stream for them until we reach a low length-storage byte.\n                \n\n#### Byte Map\n\n                  I was going to create one of these, but then I realized that unlike UTF-8000, UTF-∞-8 reuses bytes depending on context. For example 0xB4 can be\n                  `10110100`\n                  or\n                  `10110100`, and 0xAX can be\n                  `1010xxxx`\n                  or\n                  `1010xxxx`.\n                \n\n#### `strcmp(3)` Ordering\n\n                \n                  The choice of 13 byte code units being limited to 71 bits leads to a second byte of the form\n                  `100xxxxx`. This, and the choice of low length-storage bytes being of the form `1010xxxx`, high length-storage bytes being `10110100`, and the use of a unary-code-like sequence of high length-storage bytes for self-punctuation, means that UTF-∞-8 preserves `strcmp` ordering.\n                \n\n                  It seems that any `1011xxxx` 0xBX byte could have been used for the high length-storage bytes, and 0xB4 before\n\n is just a fun choice.\n                \n\n#### BOM Support\n\nFor the same reasons as UTF-8000, BOM support is maintained.\n\n### Verdict\n\nIt works, but it's quite complicated. It took me an entire day to figure out how it works, and to calculate its stats. I much prefer the simplicity of UTF-8000.\n\n                  The high length-storage bytes `10110100` provide self-punctuation and `strcmp` support, but somehow feel wasteful, taking up eight bits each, and are exceptional, with none of the other 0xBX bytes being used in a similar way. That being said, asymptotically UTF-∞-8 has a better information rate than UTF-8000.\n                \n\nThe mechanism of subtracting 18 from the number of hex digits and stuffing them into the low length-storage bytes reminds me a little of UTF-16, subtracting 0x10000 from the codepoint value and stuffing that into the surrogate bytes.\n\n## Owl's Corrected\n\n UTF-8\n\n                Owl suggests a corrected\n\n version of UTF-8 with many radical changes. Many of these are opinionated, such as removing most control codes from C0. Many are technical, such as precluding the concept of overlong encodings similarly to UTF-16, by an `n`-byte code unit decoding to the binary number stored in the content bits added to the upper bound of codepoint values from `n-1`-byte code units.\n\nEven notwithstanding the established dominance of UTF-8, I still disagree with almost everything in the document. But, he does come the closest to discovering the structure of UTF-8000's code units.\n\n### Source\n\n                  Zachary Weinberg on Owl's Portfolio: *Corrected UTF-8*\n                \n\n### TLDR / Examples\n\n| ... |  |  |  |  |  | \n| 6 | `1111110x` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` | \n| 7 | `11111110` | `10xxxxxx` | `10xxxxxx` | ... | `10xxxxxx` | \n| 8 | `11111111` | `110xxxxx` | `10xxxxxx` | ... | `10xxxxxx` | \n| 9 | `11111111` | `1110xxxx` | `10xxxxxx` | ... | `10xxxxxx` | \n| ...? |  |  |  |  |  | \n\n                  If Owl had specified the second-highest bit of his continuation start bytes to be a `0` instead of a `1` then he would have beat me to UTF-8000! So close, but so far.\n                \n\n### Properties\n\n#### No Self-Synchronization\n\n                  In the 8-byte code unit example, there is no way to distinguish the second byte `110xxxxx` from the first byte of a 2-byte code unit `110xxxxx`. This generalizes beyond just 8-byte code units. This specific example could also encode `11000001` (0xC1), which UTF-8 cannot, which may trip UTF-8 compatibility stress-tests.\n                \n\n#### BOM Collision\n\n                  Owl acknowledges that his extension may lead to issues with the UTF-16 BOM, as (presumably?) his extension permits `11111111` `11111110` (0xFF 0xFE), the little-endian UTF-16 BOM.\n                \n\n### Verdict\n\nOwl acknowledges leaving that extension for the future\n\n with respect to going beyond the 6-byte old RFC 2044 version of UTF-8, showing humility and acknowledging his design's flaws.\n\nI don't wish to dunk on his document too hard, but I'm greatly relieved that he failed to derive the infinite extension mechanism. It's not just for my ego's sake, but because I do not wish for UTF-8(000) to be associated with all the other junk in his specification.\n\n## Do Nothing\n\nWhy bother publishing this now and making so much noise? As of Unicode Version 17.0, September 9th 2025, only 299,448 of 1,114,112 (27%) codepoints have been designated.\n\n                  We choose to go to the Moon in this decade and do the other things, not because they are easy, but because they are hard, because that goal will serve to organize and measure the best of our energies and skills, because that challenge is one that we are willing to accept, one we are unwilling to postpone, and one we intend to win...\n\n\n- It is a great exercise in coding theory.\n- Nobody else seems to have figured it out, as only worse rejected alternatives have been previously proposed.\n- If we wait until we run out of codepoints, one of those rejected alternatives may be hastily implemented just because it already exists. Granted, compatibility with UTF-16 would be broken, and first codepoints with five and six byte UTF-8 representations as per RFC 2044 could be satisfactory without needing UTF-8000.\n- The current upper bound of `U+10FFFF` on codepoints is entirely due to UTF-16's maximum capacity. UTF-16 and`wchar_t` is legacy Windows tech. If the future demands a larger set of codepoints then we should allow ourselves to not be held back.\n- Who doesn't like freedom, the ability to encode any integer (unsigned or signed) that we want to?\n- I feel in charge of the intellectual property to an extent, and as such I have copylefted it, rather than allowing a tyrant to (re-)discover it and publish it on their restrictive terms.\n- ~~Fortune and~~ glory. I figured this out*myself* in the era of the rise of AI. I'll settle for the credit lol.\n- I will be submitting this work to 3b1b's Summer of Mathematics Exposition 2026.\n\n### Verdict\n\nPublish.\n\n## Feedback\n\nFeedback is welcome, by email or on GitHub, if you have any improvements or questions. I plan on reaching out to people in phases to get the most UTF-proximal\n\n feedback first. Selected feedback may go in this section.\n\n### Ken Thompson\n\nKen Thompson replied to my email. That's really cool! Here is the correspondence:\n\n## emails\n\n```\nDate: Jul 4, 2026, 1:13 AM\nFrom: Jay Berry <>\nTo: Ken Thompson <>\nSubject: I have extended UTF-8 infinitely!\nHi Ken,\nI thought you might be interested to see how (infinitely) far one can\npush UTF-8, without introducing any new special cases, and while\nmaintaining all properties like self-synchronization, `strcmp(3)`\nordering, n-byte multibyte code units having 5n+1 content bits, etc.\nI have put a one-page document on my\n[website](https://utf-8000.jb2170.com/) explaining the spec. The TLDR\nsection should be sufficient to see what's going on, splitting the\nself-synchronization bits from the self-punctuation bits, and allowing\nthe self-punctuation bits to roll over into continuation bytes.\nI have searched high and low on the internet to try to make sure that\nI have not *re*discovered this, that I am not unduly taking credit for\nit. It seems to be an original thought. I have also fairly analyzed a\nfew rejected alternatives but they all lose key properties.\nCan I ask: Did you or Rob Pike or anyone else working on FSS-UTF /\nUTF-8 intend for it to be *this* extensible / future-proof? You did a\nreally good job! At this rate it will still be around in many\ncenturies' time.\nHappy Fourth of July! Consider this a 250th birthday gift from Great\nBritain (if you'd not already thought of it while designing UTF-8 back\nin the 90s lol).\nThanks,\nJay Berry\n---\nDate: Jul 16, 2026, 11:29 PM\nFrom: Ken Thompson <>\nTo: Jay Berry <>\nSubject: Re: I have extended UTF-8 infinitely!\nyour first 2 extensions (5 and 6 bytes) were clearly envisioned.\nthe standard (up to 4 bytes) was created to cover the size of\nunicode. i thought any more description would be a waste of\npaper. i think your extension from 7 to 8 bytes is a little hoaky.\ni requires reading the whole string rather than \"knowing\" the\nnumber of follow on bytes. so, i think the only thing new is the\n7 byte version.\ni appreciate the mail, but i really dont think it is useful. it is\nlike replacing ipv6 with ipv50.\n---\nDate: Jul 17, 2026, 11:13 PM\nFrom: Jay Berry <>\nTo: Ken Thompson <>\nSubject: Re: I have extended UTF-8 infinitely!\nHi Ken,\nThanks for the reply!\nI agree with the 'ipv50' remark haha. Even if we exhaust the existing\n1,112,064 possible Unicode codepoints, going back to your original\n6-byte UTF-8 proposal yields over 2 billion codepoints (31 bits),\nwhich would be sufficient for a long while, without needing\ncontinuation-start bytes.\nI'm submitting UTF-8000 to the 2026 [Summer of Math\nExposition](https://some.3b1b.co/). I think it's still worth sharing\nif it inspires those interested in maths / computer science, even\nthough it may never be used in our lifetimes.\nDo you mind if I include this email chain in the feedback section? I\ndecided to first ask the creator of UTF-8 (yourself), then the authors\nof the alternatives that I've critiqued, then the general public.\nThanks,\nJay\n---\nDate: Jul 18, 2026, 5:41 AM\nFrom: Ken Thompson <>\nTo: Jay Berry <>\nSubject: Re: I have extended UTF-8 infinitely!\nyou can use the reply.\n                \n```\n              Only the zigzag signed variant requires reading the entire code unit (really the last bit of the last byte) to perform `strcmp` checking, not normal UTF-8000, but yes that's a good point that he's observed.\n\n### Rejected Alternatives Authors\n\nI have emailed the authors of the rejected alternatives that I've reviewed, to see what are their critiques of mine.\n\n## first email\n\n```\nDate: 2 Aug 2026, 16:14\nFrom: Jay Berry <>\nTo: Mashpoe          (UTF-Infinity) <>,\n    Larry Wall       (Perl utf8)    <>,\n    Tom Bishop       (UCS-X)        <>,\n    Richard Cook     (UCS-X)        <>,\n    Zachary Weinberg (Owl)          <>\nSubject: Unlimited UTF-8 | UTF-8000\nHi everyone!\nI believe that I have discovered the \"correct\" way to extend UTF-8 infinitely,\nwithout introducing any new special cases, and while maintaining all properties\nlike self-synchronization, self-punctuation, `strcmp(3)` ordering, n-byte multibyte\nunits having 5n+1 content bits, etc. I've codenamed it \"UTF-8000\" or \"UTF-8K\".\nI have put a one-page document on my [website](https://utf-8000.jb2170.com/)\nexplaining the spec. The TLDR section should be sufficient to see what's going on,\nsplitting the self-synchronization bits from the self-punctuation bits, and\nallowing the self-punctuation bits to roll over into continuation bytes.\nReference implementation in Python is available on\n[GitHub](https://github.com/UTF-8000/UTF-8000-Python) which can be installed\nwith `$ pipx install UTF-8000`.\nI noticed that each of you have attempted to extend UTF-8 in different ways,\nand I have constructively reviewed each of them in the\n[rejected alternatives](https://utf-8000.jb2170.com/#sec-rejected-alternatives)\nsection of my spec. I thought you might be interested / maybe you have some\nfeedback for mine.\n- [UTF-Infinity](https://utf-8000.jb2170.com/#sec-rejected-utf-infinity)   by Mashpoe\n- [Perl utf8](https://utf-8000.jb2170.com/#sec-rejected-perl-utf8)         by Larry Wall\n- [UCS-X](https://utf-8000.jb2170.com/#sec-rejected-ucs-x)                 by Tom Bishop and Richard Cook\n- [Owl's \"Corrected\" UTF-8](https://utf-8000.jb2170.com/#sec-rejected-owl) by Zachary Weinberg\nI emailed Ken Thompson, creator of UTF-8 (and Unix!) to see what he thinks,\nand I got a reply! The exchange is on the\n[website](https://utf-8000.jb2170.com/#sec-feedback-ken-thompson). His remark\n\"it is like replacing ipv6 with ipv50\" is funny to me and should set a\nnot-too-serious atmosphere for this whole discussion.\nNonetheless I still think that UTF-8000 is fun and educational anyways, an exercise\nin coding theory, and I'll be submitting it to 3b1b's 2026\n[Summer of Math Exposition](https://some.3b1b.co/). But first I thought that it\nwould be proper to email the people whose work I've reviewed.\nThanks,\nJay Berry\n                \n```\n              #### Zachary Weinberg (Owl)\n\n## emails\n\n````\nDate: 2 Aug 2026, 19:18\nFrom: Zachary Weinberg <>\nTo: Jay Berry <>\nSubject: Re: Unlimited UTF-8 | UTF-8000\nOn Sun, Aug 2, 2026, at 11:14 AM, Jay Berry wrote:\n> I believe that I have discovered the \"correct\" way to extend UTF-8\n> infinitely, without introducing any new special cases, and while\n> maintaining all properties like self-synchronization, self-\n> punctuation, `strcmp(3)` ordering, n-byte multibyte units having 5n+1\n> content bits, etc. I've codenamed it \"UTF-8000\" or \"UTF-8K\".\nHey, thanks for reaching out.  I'm delighted to see that I am not the\nonly one fed up with the artificial limitation of UTF-8's encoding space\nto match UTF-16. I may actually revise my proposal to adopt your trick\nfor preserving self-synchronization even when the start bits extend\npast the end of the first byte.\nI think you're not taking the value of *eliminating* overlength\nencodings seriously enough, though.  Yeah, that's the most complicated\npart of my proposal, and the part that means Corrected UTF-8 doesn't\ncorrespond to IETF UTF-8 for anything but the ASCII page, but it's also\nthe part that means Corrected UTF-8 decoders *cannot* be a vehicle for\npath-smuggling attacks on network services, and therefore I consider it\nsecond in importance only to lifting the artificial plane limit.\nMaking it impossible to encode surrogates is also important for\nsecurity reasons; the only way I could be persuaded to not do that is\nif there was any chance that the surrogates might get *reassigned* as\nordinary characters in a couple decades, once UTF-16 is truly dead and\nburied ... and I think the odds of that ever happening are far lower\nthan the odds of the Unicode Consortium backing down on their \"never will\nthere be more than 17 planes\" policy.\nI've mostly come around to agree with you on the C1 controls, though.\nOmitting them doesn't really help anything.  At the time I wrote the\noriginal document (some years before I posted it on my website) I was\nstill working at a browser company and mislabeled or mistranscoded\nWindows-1252 was a regular headache; but I have the impression that\nhas become much less common over the past decade and a half, and\n\"this can represent any Unicode codepoint with an official non-surrogate\nassignment\" *is* a desirable property for anything calling itself an UTF.\nzw\n---\nDate: 3 Aug 2026, 23:34\nFrom: Jay Berry <>\nTo: Zachary Weinberg <>\nSubject: Re: Unlimited UTF-8 | UTF-8000\nHi Zack,\nThanks for the reply!\n> I may actually revise my proposal to adopt your trick\n> for preserving self-synchronization even when the start bits extend\n> past the end of the first byte.\nSelf-synchronization was indeed a main feature that Ken Thompson figured out\nin fixing FSS-UTF:\n```\n0vvvvvvv\n10vvvvvv 1vvvvvvv\n110vvvvv 1vvvvvvv 1vvvvvvv\n...\n```\n(in which one couldn't tell the difference between eg a 2-byte start byte\n`10|vvvvvv` and a continuation byte `1|0vvvvvv`) to UTF-8:\n```\n0vvvvvvv\n110vvvvv 10vvvvvv\n1110vvvv 10vvvvvv 10vvvvvv\n...\n```\nin which the self-synchronization prefixes `0`, `10`, and `11` are distinct.\n[History of FSS-UTF -> UTF-8](https://www.cl.cam.ac.uk/~mgk25/ucs/utf-8-history.txt\n#:~:text=10zzzzzz%201yyyyyyy). It is definitely well worth keeping :)\n> I think you're not taking the value of *eliminating* overlength\n> encodings seriously enough, though.\nHaving n-byte (modified) UTF-8 decode to (value of content bits) plus\n(1 more than maximum that (n-1)-byte UTF-8 can encode) i.e. the offsets in\nyour specification, sounds good at first, but if n is big then the offset accrues:\n2^7 + 2^11 + 2^16 + ... + 2^(5(n-1)+1). We can write this as\n2^7 + 2^11 * ((2^5)^0 + (2^5)^1 + ... + (2^5)^(n-3)) as a geometric series and\nexplicitly compute it as 2^7 + 2^11 * (32^(n-2) - 1) / 31 for n >= 3,\nbut that division is a bit 'icky' compared to addition subtraction multiplication\nand bitshifting.\n```py\ndef offset(n: int) -> int:\n    if n == 1:\n        return 0\n    elif n == 2:\n        return (1 << 7)\n    else:\n        return (1 << 7) + (1 << 11) * ((1 << (5 * (n - 2))) // 31)\n```\nIt seems a lot easier to say \"n-byte multibyte UTF-8 can store up to\n(5n+1)-bit codepoints\", a nice instance being 3-byte UTF-8 storing 16 bits,\n1 2 and 3 byte UTF-8 exactly covering\n[Plane 0](https://en.wikipedia.org/wiki/Plane_(Unicode)) of Unicode.\n[UTF-1](https://en.wikipedia.org/wiki/UTF-1) was an earlier encoding that\nKen Thompson and Rob Pike tried out\n([interview](https://www.youtube.com/watch?v=OmVHkL0IWk4&t=14275s)). It used\n`mod 190` and divisions which they disliked, and eventually FSS-UTF and UTF-8\ncame around which use simple bitwise operations to check against overlong encodings,\nand to extract the content bits. I think the anti-overlong checking is not too\ncomplicated, 2-byte UTF-8 being the only odd one out.\nUnicode offered a way forward from the ISO 8859-{1..16} diaspora of 8-bit codepages.\nUTF-8 offered ASCII forwards compatibility with better efficiency than UTF-16 using\nbyte-precision rather than word-precision. I don't think that your offset-based\nUTF-8 offers a significant upgrade. It's really just to protect noob software\ndevelopers who might write broken decoders for UTF-8 that don't do anti-overlong\nchecking and surrogate range checking.\n> any chance that the surrogates might get *reassigned* as\n> ordinary characters in a couple decades\nMy guess is that regardless of whether UTF-16 continues to live on, the\nsurrogate range will remain unencodable, due to pre-established UTF-8 parsers\nrejecting them. Though I could be wrong, as iirc IP addresses that ended in `.0`\nwere originally not allowed, but now are. It does make one wonder what those\n2048 codepoints could be assigned to...\n> I've mostly come around to agree with you on the C1 controls, though.\nThat's great. I don't think I've ever actively used them, but yeah C1 should be\nin / remain in Unicode as a way to refer to it using codepoints. It's a bit of a\nshame that they don't (yet) have a 'control pictures' block like\n[C0 Does](https://www.compart.com/en/unicode/block/U+2400).\n> a regular headache\nOn a similar note there's still some software that is fiddly with them. I've got\na [bug to fix](https://github.com/jb2170/better-adb-sync/issues/42) that I've\nfigured out the solution to whilst messing with UTF-8(000) and encodings.\nAndroid's Toybox's `ls` outputs U+0080 to U+009F 'C1 control codes' and\nU+00A0 'non-breaking space' differently depending on whether `ls` is running over\n`adb` interactively or not: interactively U+00A0 prints as `\\240` octal-escape style,\nbut non-interactively it prints as the byte `a0`, which is either a coincidentally\ndecapitated UTF-8 unit `c2 a0`, or the raw ISO-8859-1 8-bit byte. The fix is to\nuse the `-b` flag on `ls` to force an escaped style. So tldr I understand the\nfiddly-ness lol.\nThanks,\nJay\n                \n````\n              #### Tom Bishop (UCS-X)\n\n## emails\n\n```\nDate: 2 Aug 2026, 19:41\nFrom: Thomas Eugene Bishop <>\nTo: All <>\nSubject: Re: Unlimited UTF-8 | UTF-8000\nHi Jay,\nThanks for letting me know about your work, and the others you reference. It's\ngood to know that others are interested in extending the range of encoding.\nI'll study the proposals more when I have time. Based on first impressions, I\nhave these comments.\nAbout \"the 'correct' way\": maybe you mean that ironically and recognize there's\nmore than one way to do it, with trade-offs. On the other hand, you wrote,\n\"Nobody else seems to have figured it out, as only worse rejected alternatives\nhave been previously proposed.\" That sounds like an unwarranted claim that\nyou've solved a problem nobody else was able to solve. I wish you wouldn't use\nthe word \"rejected\" to describe alternatives, since it might be misconstrued\n(maybe through an AI search) as implying a decision by an organization with some\ncapacity to accept or reject proposals. I think what you mean is that you\npersonally prefer your own proposal.\nNow that multiple solutions exist, there's room to compare them by various\ncriteria such as efficiency, simplicity, and robustness.\nYou described Larry Wall's utf8 as \"inextensible\"; that's wrong, as proved by\nits extension to UTF-∞-8. Or else, \"inextensible\" doesn't mean what I think it\nmeans. Also, my understanding is that the contrast between \"utf8\" and \"UTF-8\"\nwas intentional.\nYou wrote, \"UCS-X proposes three extensions for each of UTF-8, UTF-16, UTF-32,\nfor a total of nine specifications, twelve including the existing base\nspecifications!\" and \"it's quite complicated\". I think this reference to 12\nspecs is an unfair criticism. The existence of multiple specs doesn't imply\ncomplexity of the encodings themselves. The complication of the existing 3 base\nspecs is obviously beyond anybody's control at this point. Of the remaining 9,\nyou can ignore 6 if you want, since they are merely simplifications of the last\n3; that is, the specs with max U+7FFFFFFF and U+7FFFFFFFFFFFFFFF are just\nsubsets of the specs with max infinity. We separated them out to support\nimplementers who might have good reasons not to go straight to infinity.\nYou wrote, \"At a glance the UTF-16 extensions look like they break syntax with\nbase UTF-16, ...\". I don't know what you mean by \"break syntax\", but UTF-∞-16 is\na compatible extension of UTF-16 in the sense that our spec defines \"compatible\nextension\". It would be more responsible to postpone publishing a \"break syntax\"\nassertion until you're certain and ready to explain what you mean by it.\nYou wrote, \"... UTF-8000 code units can be arbitrarily large\" -- I think you\nmean UTF-8000 codes can be arbitrarily large. A UTF-8000 code unit is always 8\nbits, right?\nTo me, while the technical details of encoding are interesting, what's more\ninteresting is how people might eventually use extended encodings, such as to\ndefine their own characters and use them for public communication, without\nhaving to wait for official approval of each character.\nBest wishes,\nTom\n---\nDate: 3 Aug 2026, 16:57\nFrom: Thomas Eugene Bishop <>\nTo: All <>\nSubject: Re: Unlimited UTF-8 | UTF-8000\nHi Jay,\nI wrote a script to compare the lengths of UTF-8000 and UTF-∞-8 codes, and also\ntheir \"start\" bytes. That script isn't thoroughly tested and it might be only\napproximate especially in some edge cases. With that disclaimer, it seems that\nif a USV has 47 or more digits, then UTF-8000 is longer than UTF-∞-8. If a USV\nhas 18 or more digits, the number of \"start bytes\" (needed to determine the\nlength of an entire code) is longer for UTF-8000 than for UTF-∞-8. For a USV\nwith 128 digits, UTF-8000 has 102 total bytes and 17 start bytes, while UTF-∞-8\nhas 90 total bytes and 4 start bytes. UTF-8000 does have shorter codes in some\nranges, such as for USV with 10-15 digits.\nNeither solution is optimal in terms of storage size. There are trade-offs such\nas speed of execution, simplicity, robustness, etc.\nThe number of start bytes might be important in situations where text is read\ninto a fixed-size buffer and a buffer might contain a partial code. Software\nshould be able to determine the length of an entire code by scanning a\nrelatively small number of start bytes, both for efficiency and to avoid bugs in\ncases where one code might span many buffers. This is an example of\n\"robustness\". Another example is that protocols should enable processes to\nindicate max supported USV.\nThe term \"code unit\" has a standard definition\n(https://unicode.org/glossary/#code_unit) that differs from yours\n(https://utf-8000.jb2170.com/#def-code-unit). I recommend following the standard\nto avoid confusion.\nIt's wonderful that you might bring up this topic at the Summer of Math Exposition!\nCheers,\nTom\n---\nDate: 8 Aug 2026, 21:26\nFrom: Jay Berry <>\nTo: Thomas Eugene Bishop <>\nSubject: Re: Unlimited UTF-8 | UTF-8000\nHi Tom,\nThanks for the feedback!\n> About \"the 'correct' way\": maybe you mean that ironically and recognize\n> there's more than one way to do it, with trade-offs.\nThere are indeed other solutions such as UTF-∞-8 which preserve all properties\nlike self-synchronization, self-punctuation, strcmp order etc. However the\nreason I've referred to it as the 'correct' way is because in my opinion it\nlooks like the 'natural' way to extend UTF-8, as I put in the [properties]\nsection addressing the fact that the anti-overlong mechanism works the same as\nUTF-8, with no new special cases. I find it very simple to explain (in\nretrospect) to begin with bytes endowed with self-synchronization prefixes '11'\nand '10', and to stripe the self-punctuation bits across them.\n> you wrote, \"Nobody else seems to have figured it out, as only worse rejected\n> alternatives have been previously proposed.\"\nBy that I mean that nobody else online has suggested the exact layout that\nUTF-8000 proposes, which I feel is the 'natural' / 'correct' one, formally\nidentifying the self-synchronization and self-punctuation bits and how to use\nthem. The verdicts on the other proposals summarize their flaws, all but UTF-∞-8\nlosing key properties of interest.\n> I wish you wouldn't use the word \"rejected\" ... implying a decision by an\n> organization with some capacity to accept or reject proposals\nI styled my document a bit like a [Python PEP], in which often the alternatives\nhave to be firmly disproven. That being said, yes I don't think I've made it\nclear that this is a *proposal*, not an existing standard. In the Python\nreference implementation [readme] I added the line \"UTF-8000 is in no way\nendorsed by or representative of the Unicode Consortium. This is a standalone\nproject.\". I think I'll copy that to the header of the website, thanks!\n> You described Larry Wall's utf8 as \"inextensible\"\nI know it looks like I'm contradicting myself 10 seconds later by pointing out\nthat UCS-X extends from utf8, but what I meant is that Perl utf8 *on its own* is\ndesigned only to go up to 2^63-1. It uses the `FF` byte to start its 13-byte\nunits and doesn't specify how one could continue onwards. I also don't feel the\nneed for UTF-8000 to extend utf8 like UCS-X does, as utf8 is not used outside\nPerl, and we have the opportunity to make UTF-8000 more flexible allowing\n8,9,10,11,12 byte units with 'correct' self-punctuation syntax (whereas utf8's\nsecond byte is just a plain 0x80).\n> Also, my understanding is that the contrast between \"utf8\" and \"UTF-8\" was intentional.\nYeah I'll remove that line about \"utf8\" vs \"UTF-8\", thanks. I know that the\nUnicode Consortium is pedantic with referring to 'UTF-8' using a hyphen, and I\noriginally thought that Perl was just being a bit loose with the naming. It is\nmore likely that 'utf8' was chosen to show that it's not *exactly* 'UTF-8', like\nI'm using 'UTF-8000' as a codename for my proposal.\n> I think this reference to 12 specs is an unfair criticism.  The existence of\n> multiple specs doesn't imply complexity of the encodings themselves.  We\n> separated them out to support implementers who might have good reasons not to\n> go straight to infinity.\nWe can group UTF-8 and UTF-G-8 together since they both follow the same style,\nand 5/6-byte UTF-8 was envisioned by Ken Thompson. As for the UTF-E-8 and\nUTF-∞-8 specifications, they are very different.\nI think that UTF-8000, which is just one specification, within which there are\n'natural ranges' (ie limiting to n-byte units) is a better approach. My original\nspecification for UTF-16K was going to use just one Unicode Plane, to provide\ndecent efficiency but without being too greedy in needing to claim existing\nUnicode codepoints. But then I realised that this would provide 14k+1 content\nbits, whereas if we used two planes instead of one, this would be 15k+1 content\nbits, which overlaps nicely with 5n+1 provided by UTF-8000. So I have taken some\nthought and care as to create 'ranges' like your 'Giga', 'Exa', 'Inf' ideas,\nwith which UTF-8 and UTF-16 can be expanded in parallel. I put this in the\n[UTF-16K spec]. I think it's a lot easier to say \"this is what n-byte UTF-8 and\nk-surrogate-pair UTF-16 looks like. restrict to n=3k and you have ranges that\nencode the same codepoints\" than to have a patchwork of different standards\nbased on what range a codepoint is in, like UCS-X eg includes Perl utf8 as\nUTF-E-8.\n> I don't know what you mean by \"break syntax\", but UTF-∞-16 is a compatible extension\nBy 'compatible' I'm thinking along the lines of backwards compatibility \"will\nthis throw an error in a UTF-8 / UTF-16 parser?\" and \"are we maintaining the\npre-established syntax?\".\nFor UTF-8, technically one could argue that UTF-8000 and UTF-∞-8 \"break syntax\"\nby eg using the byte 'FF', which when fed into a UTF-8 parser will cause an\nexception. However on the other hand the byte 'FF' causing an exception is only\ndue to the restriction to U+10FFFF on the range of codepoints for UTF-8,\nprovided one's UTF-8 extension uses the byte 'FF'. So yes I'm being a bit\nhypocritical, but I feel fine with that because bytes F{5..F} are currently\nunused by UTF-8, and the proposed syntax of UTF-8000 is predictably the same as\nUTF-8, eg wrt self-synchronization prefixes for non-ASCII first bytes being\n'11', and for continuation bytes being '10'.\nFor UTF-16, every 16-bit word has already been used. Instead of changing the\nsyntax to use eg 1 high surrogate and (n-1) low surrogates, or like UTF-G-16 use\nn low surrogates, I decided to use a \"semantic reinterpretation\" layer on top of\nUTF-16, ASCVI-on-UTF-16. Ie, just as UTF-16 is a semantic reinterpretation of\nUCS-2, interpreting codepoints in the ranges U+D800 to U+DBFF and U+DC00 to\nU+DFFF no longer as those individual codepoint values, but rather as parts of\nsurrogate pairs, so too I decided to implement UTF-16K as a semantic\nreinterpretation of plane 9 and 10 surrogate pairs. The nice thing about this is\nthat a decoder which only understands UTF-16 can open a UTF-16K encoded file,\njust as a UCS-2 decoder can open UTF-16 files. Plane 9 and 10 surrogate pairs\nwould be displayed as UTF-16 codepoints rather than as one UTF-16K codepoint,\njust as a UCS-2 parser would show two surrogate codepoints instead of one UTF-16\ncodepoint; semantic errors rather than syntax errors. Contrast that with\nUTF-G-16, U+110000 encoded as 'DC04 DE80 DE00', with which the opening word may\nimmediately raise an exception in a UTF-16 parser.\nFor UTF-G-16, for ill-formed units, I am able to generate context-dependent\nerror handling behavior which leads to errors being decoded as though they are\ncorrect. I am able to cause a contradiction in your UTF-G-16 [decoding rules]:\nmake 'DC04' both preceded by D800 (to make it trailing) and succeeded by DE80\n(to make it leading). If we were to seek to the point 'X' in a stream 'X D800 Y\nDC04 DE80 DE00' we would decode this as 'U+10004 (D800 DC04) U+FFFD (replace\nDE80) U+FFFD (replace DE00)'. If we were to seek to the point 'Y' we would\ndecode this as 'U+110000 (DC04 DE80 DE00)', using the low surrogate 'DC04' and\nleaving the high surrogate 'D800' before the seek point Y. This looks like bad\nbehavior. In UTF-8 and UTF-8000 because the first-byte and continuation-byte\nself-synchronization prefixes make their byte ranges disjoint, I don't think a\nsituation like this can happen there. Ie never will a 'well formed unit X\nfollowed by errors' be incorrectly decoded as a 'well formed unit Y with perhaps\nsome junk before it' if one seeks to the middle of the well formed unit 'X'. So\ntoo UTF-16K keeps the {high surrogate | low surrogate} and {first surrogate pair\n(plane 9) | continuation surrogate pair (plane 10)} ranges disjoint which avoids\nthis issue and maintains self-synchronization at the word-level. UTF-G-16\nmuddies the water by 'DC04' being trailing (UTF-16 surrogate pair) or leading\n(UTF-G-16 leading) dependent on previous words. This is also why a UTF-8 /\nUTF-8000 parser only ever needs to seek *forwards* to the next first byte if it\nencounters an error.\n~~For UTF-G-16, for well formed units, something still doesn't feel right that\none might need to look backwards to determine whether eg 'DC04' is trailing or\nleading. We do not always have backwards seeking, like on a pipe or socket, or\nat least we don't want to do backtracking like complicated regexes sometimes\ndo.~~ In well formed units we know exactly one of those conditions will be true\nand we can look forwards rather than back, right? This seems like minutiae\ncompared to the behaviour in the previous paragraph.\nBack to UTF-8, this conversation has made me realize that one could implement\n*private-use extensions* on top of Unicode / UTF-8 using ASCVI-on-UTF-8, in a\nsimilar way to UTF-16K using ASCVI-on-UTF-16. We can achieve an ASCVI-like code\nin as little as 3 bits, 8 codepoints:\n0: 000, 1: 001, 2: 010 110, 3: 010 111, 4: 011 101 100, 5: 011 101 101,\n6: 011 101 110, 7: 011 101 111, 8: 011 110 110 100, 9: 011 110 110 101, ...\nthough using more bits will of course lead to more efficient codes. The\nadvantage of this style is that it's just a semantic reinterpretation layer on\ntop of UTF-8, and will pass right through a UTF-8 parser okay. A good range of\ncodepoints to use may be some of the U+E000 to U+F8FF Plane 0 private-use\ncodepoints. This seems like a great way in which one could create their own\nautonomous set of 'MyUnicode' codepoints M+...XXXX starting at M+0000,\nMyUnicode-on-Unicode style (as opposed to UTF-16K which uses the *public*\nUnicode range and postulates starting at U+110000). This would answer your\npoint:\n> what's more interesting is how people might eventually use extended encodings,\n> such as to define their own characters and use them for public communication,\n> without having to wait for official approval of each character\nIt does somewhat go against the spirit of \"Uni\"code, which is the one-and-only\n'flat' layer of codepoints, to use an ASCVI layer on top of Unicode / UTF-8. One\ncan also imagine ASCVI-on-(ASCVI-on-UTF-8) if the M+...XXXX codepoints had their\n*own* private-use area which allowed further sub-encoding. It's a fun thought to\nthink of trees of Unicode embedded recursively as layers on top of each other,\nbut it would surely be a bit anarchic and low-efficiency. Therefore my main\nfocus with UTF-8000 and UTF-16K has been on how *Unicode* could expand in the\nlong run. The private-use extensions do sound fun, but may be a bit clunky when\ndecoded in a programming language, being interspersed in 'normal' Unicode\nstrings.\n> I wrote a script to compare the lengths of UTF-8000 and UTF-∞-8 codes, and\n> also their \"start\" bytes.\nYes UTF-∞-8 has shorter units in the long run, an efficiency tending towards 6/8\nwhereas UTF-8000's efficiency tends towards 5/8. It's probably easiest to point\nto UTF-8000 using a linear number of self-punctuation bits (n-1) -> O(n),\nwhereas UTF-∞-8 is roughly logarithmic O(log_2(n)).\n> Software should be able to determine the length of an entire code by scanning\n> a relatively small number of start bytes, both for efficiency and to avoid\n> bugs in cases where one code might span many buffers.\nThis is a good point, and UTF-∞-8 is more succinct with respect to\nself-punctuation. For 33 hex-digit codepoints, UTF-∞-8 uses 2 bytes, whereas\nUTF-8000 uses 5, for 273 hex-digits UTF-∞-8 uses 4 bytes, whereas UTF-8000 uses\n37! Mogs me.\nand finally\n> The term \"code unit\" has a standard definition that differs from yours. I\n> recommend following the standard to avoid confusion.\nI realised this half way through writing the UTF-8000 spec and I'm struggling to\nthink of an alternative name. I opened a GitHub [issue] to remind me to rename\nit. 😅\nSo to conclude so far:\n- I still think that UTF-8000 is simpler to explain and more predictable than UTF-∞-8\n- UTF-∞-8 is asymptotically more efficient than UTF-8000 and requires less start\n  bytes (self-punctuation bytes)\n- UTF-G-16 (and beyond?) looks broken to me, though I haven't properly anatomized\n  the UTF-X-16 family of UCS-X proposals like I have for the UTF-X-8 family.\n- ASCVI-on-private-use-UTF-8 sounds like an okay idea for private-use extensions\n  if they require a large amount of 'codepoints' (sub-encoded virtual\n  my-codepoints M+...XXXX)\n- I have a few remarks to change on my proposal\nThis has been a fun project! Thanks for the emails,\nJay\n                \n```\n              I have updated the UTF-16K specification to mention UCS-X's UTF-G-16's flawed error handling behavior.\n\n### SoME 2026\n\nI am submitting this work to 3b1b's Summer of Mathematics Exposition 2026. I hope that it is useful to some people who view it, that it is educational about coding theory, and that maybe there'll be some feedback.\n\n### General\n\nI'll put a link to this page on r/Unicode. There's lots of show-and-tell on there.\n\n### Unicode Consortium\n\nI *might* send this to the Unicode Consortium if there is good consensus from the feedback above.\n\nBut as Ken put it we don't really need IPv50 or unlimited UTF-8 right now, so I don't want to pester the Unicode Consortium when they're busy doing actually important jobs like documenting scripts, assigning codepoints, helping internationalization, etc.\n\nMaybe this specification can sit in a 250-year time capsule in the Unicode Consortium Archives for when the time is right to expand...\n\n## Reference Implementation and Tools\n\n### UTF-8000\n\nWorking reference implementation in Python with comprehensive code documentation is available on GitHub as UTF-8000/UTF-8000-Python. It can be installed as a PyPI package using `$ pipx install UTF-8000` which provides the command line utility `utf-8000(1)`.\n\nThe `$ utf-8000 info` subcommand displays info about a codepoint encoded in UTF-8000, with useful bit highlighting.\n\nThe `$ utf-8000 encode` subcommand reads codepoints from stdin and writes the raw UTF-8000 bytes to stdout.\n\nThe `$ utf-8000 decode` subcommand reads UTF-8000 bytes from stdin and feeds them to an incremental decoder, writing the decoded codepoints to stdout.\n\n### UTF-16K\n\nWorking reference implementation for UTF-16K is also available on GitHub as UTF-8000/UTF-16K-Python. It can be installed using `$ pipx install UTF-16K` which provides `utf-16k(1)` with the same subcommands as `utf-8000(1)`.\n\n## Naming\n\nIn the development phase of this project I have been using the codename UTF-8000\n\n, but I find myself increasingly drawn to UTF-8K\n\n.\n\nBelow is a comparison of different potential names, and I am open to suggestions.\n\n### UTF-8000\n\nInspired by Python 3's development codenames in PEP 3000.\n\n#### Pros\n\n- The thousand inUTF eight thousand sounds big and futuristic. This encoding scheme should also last forever!\n\n#### Cons\n\n- The zeros are repetitive.\n- Having to remember exactly three zeros to make 8000 . Some people may read it as eight hundred, or eighty thousand etc.\n- The inspiration logic doesn't exactly match up with Python because Python was moving from Python 2 to Python 3, not Python 3 to Python 3000, whereas we're going from UTF-8 to UTF-8000.\n- UTF-8 is a prefix ofUTF-8000 . Existing software which parses a string representing the encoding's name to determine the encoding may do something like`if encoding_name[:5] == \"UTF-8\"` and incorrectly short-circuit. It is for this reason that Microsoft skipped from Windows 8 to Windows 10, not creating Windows 9, because existing software may check forWindows 9 to test forWindows 95 orWindows 98 .\n\n### UTF-8K\n\nInspired by another of Python 3's development codenames Py3K\n\n / Py3k\n\n, we could use UTF-8K, the K\n\n capitalized like the UTF\n\n. It's also shorter than 8000\n\n.\n\n#### Pros\n\n- Short; just one more letter than UTF-8 .\n\n#### Cons\n\n- UTF-8 is a prefix ofUTF-8K . See here.\n\n### UTF-8 & Knuckles\n\nThe K\n\n in UTF-8K\n\n reminds me of Sonic 3 & Knuckles\n\n, sometimes abbreviated to S3K\n\n.\n\n                The Sonic & Knuckles\n\n cartridge uses lock-on technology\n\n to extend Sonic 3\n\n and Sonic & Knuckles\n\n into Sonic 3 & Knuckles\n\n. In a similar way UTF-8 extends from (locks on to) ASCII, and UTF-8000 continues this extension.\n              \n\n#### Pros\n\n- Knuckles is cool.\n\n#### Cons\n\n- SEGA might not be happy, though they do seem nicer than Nintendo with respect to fanart.\n- The hardest work of extending from ASCII to UTF-8 has already been achieved by Ken Thompson and Rob Pike. Is UTF-8000 lock-on if it doesn't introduce any new special cases? Not really.\n- UTF-8 is a prefix ofUTF-8 & Knuckles . See here.\n- This is just a bit of a joke-y name.\n\n### VTF-8\n\nShort for Variable Transformation Format 8\n\n.\n\n#### Pros\n\n- Removes the association with Unicode as VTF-8 is just a method of storing unsigned integers.\n- UTF-8 is not a prefix ofVTF-8 ; in fact they don't even begin with the same letter.\n- The V looks Romanesque and fits with the logo.\n\n#### Cons\n\n- The V inVTF-8 looks very similar to theU in UTF-8. At a glance, and depending on font rendering, one may not notice the difference.\n\n### STF-8 for the Signed Variant\n\nThe U\n\n in UTF-8\n\n could also be read as Unsigned\n\n, ie Unsigned Transformation Format 8\n\n. Thus we could take inspiration and write STF-8\n\n for Signed Transformation Format 8\n\n.\n\n## Logo\n\nUTF-8 does not have a logo. I had a bit of fun designing a logo for UTF-8000. I made it Romanesque but simple.\n\nIt's an eagle with UTF\n\n written across its wings and chest. The Roman numerals for 8\n\n, VIII, flank its head.\n\n                Below its claws it holds a fasces, wrapped with a continuation byte of the form\n                `10xxxxxx`. This represents the united strength of a bundle of continuation bytes, which makes us unstoppable in conquering the entire integers, in the name of including them into Unicode codepoints, encoded in UTF-8000.\n              \n\nThe central fascis which sticks out represents the first byte of a code unit. This byte looks different from its continuation bytes, and has the power to start a code unit, which is represented by the wielding of the axe, sticking out on the left. On the right end the central fascis sticks out perhaps representing the terminating 0 of the start bit sequence.\n\nThe favicon for this website is just the Roman numerals VIII.\n\nThe Git repo containing these images is available on GitHub as UTF-8000/UTF-8000-Images.\n\n## Licensing / Copyright (Copyleft)\n\nAs creator of UTF-8000 I want the liberty with which UTF-8000 (the algorithm) can be used to be no less permissible than UTF-8 and ASCII before it. This belongs to everyone. Live Free or Die.\n\nThis website is licensed under CC-BY-NC-SA-4.0, available on GitHub as UTF-8000/UTF-8000-Website.\n\nThe logo / images are licensed under CC-BY-NC-SA-4.0, available on GitHub as UTF-8000/UTF-8000-Images.\n\nThe Python reference implementation is licensed under GPL-3.0-only, available on GitHub as UTF-8000/UTF-8000-Python. Any implementation of decoding and encoding UTF-8000 is going to look somewhat similar to this codebase of course; don't worry if you want to use MIT or BSD or something else in a clean-room rewrite.\n\n## Thanks\n\nBell Labs:\n\n- \n                  Claude Shannon, founding father of the Digital Age, Information Theory, and Artificial Intelligence. His 1948 paper *A Mathematical Theory of Communication* is the most important mathematics paper of the mid 20th Century, which forms a good chunk of the University of Cambridge Mathematics Tripos course*Coding and Cryptography* . I made a Manim YouTube video in 2024 covering Entropy from Information Theory and its various appearances in mathematics.\n- Ken Thompson, creator of UTF-8, also best known for Unix and other big projects like chess computers.\n- Rob Pike, co-creator of UTF-8, also best known for Plan 9 and other big projects.\n\nThe University of Cambridge mathematics department:\n\n- Professor Stuart Martin, who lectured *Coding and Cryptography* 2019-2020.\n- \n                  Dr Ross Lawther, director of studies for mathematics at Girton College who supervised me for *Coding and Cryptography* (and other mathematics topics!) 2017-2020. It is question 2 of example sheet 1 that concerns the product of two prefix-free codes, which I was reminded of by the product of the self-synchronization and self-punctuation mechanisms of UTF-8000.\n- Dr Keith Carne, whose timeless lecture notes for *Codes and Cryptography* I use often. They are mirrored on my website here. Start at chapter 3 if you're interested!\n\n## Epilogue\n\n                To think of these stars that you see overhead at night, these vast worlds which we can never reach. I would annex the planets if I could.\n\n- Cecil John Rhodes, founder of Rhodesia.\n\n\nAnd annex the entire integers we have done! I find it fitting that ASCII (🇺🇸) and UTF-8 (🇺🇸) are completed by UTF-8000 (🇬🇧), another Anglosphere classic. But I do have another idea:\n\nI am publishing this document formally on July 4th 2026, perhaps as a 250th birthday gift from Great Britain to the United States of America. Cheers!\n\n                One man alone in a room with a computer, a typewriter as it was, can change the world.\n\n- Jonathan Bowden, English cultural orator.\n\n(As a proud alumnus of Girton College I do feel compelled to comment that man should be read in the Lockean sense of mankind!)\n\n\nThe world has been changed at least twice by Ken Thompson, with Unix and UTF-8. Unix was first written in solitude, in three weeks of the summer of 1969 (the same time as the Moon landing), at Bell Labs on a teletypewriter attached to a spare PDP computer. UTF-8 was invented in one night of autumn 1992, at a New Jersey diner on a placemat, with a slight tweak a few days later.\n\nI, Jay Berry, have so far written this document alone in the summer of 2026, realizing that my reference implementation from autumn 2024 was a nontrivial discovery.","body_html":"<h1 id=\"utf-8000\">UTF-8000</h1>\n<p>Unlimited UTF-8! ASCII ⊆ UTF-8 ⊆ UTF-8000.</p>\n<p>No special cases introduced. All properties preserved.</p>\n<p>Try out the reference implementation with <code>$ pipx install UTF-8000</code>.</p>\n<pre><code>            UTF-8000 is in no way endorsed by or representative of the Unicode Consortium. \n\n            This is a fun standalone project / proposal.</code></pre>\n<h2 id=\"tldr-examples\">TLDR / Examples</h2>\n<p>| ASCII |  |  |  |  |  |  |  |  |  |  |  | \n| 1 | <code>0xxxxxxx</code> |  |  |  |  |  |  |  |  |  |  | \n| UTF-8 |  |  |  |  |  |  |  |  |  |  |  | \n| 2 | <code>110xxxxx</code> | <code>10xxxxxx</code> |  |  |  |  |  |  |  |  |  | \n| 3 | <code>1110xxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  |  |  |  |  |  |  | \n| 4 | <code>11110xxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  |  |  |  |  |  | \n| UTF-8000 |  |  |  |  |  |  |  |  |  |  |  | \n| 5 | <code>111110xx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  |  |  |  |  | \n| 6 | <code>1111110x</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  |  |  |  | \n| 7 | <code>11111110</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  |  |  | \n| 8 | <code>11111111</code> | <code>100xxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  |  | \n| 9 | <code>11111111</code> | <code>1010xxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  | \n| 10 | <code>11111111</code> | <code>10110xxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  | \n| ... |  |  |  |  |  |  |  |  |  |  |  | \n| 22 | <code>11111111</code> | <code>10111111</code> | <code>10111111</code> | <code>10110xxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | ... | <code>10xxxxxx</code> | \n| ... |  |  |  |  |  |  |  |  |  |  |  | </p>\n<p>There is nothing special-case-y about the example 22-byte code unit here. It is just a good prototypical example, demonstrating the power of UTF-8000 with multiple start bytes.</p>\n<p>There are only two special cases, both of which are inherited from UTF-8: ASCII as is, and 2-byte UTF-8 having 4 mandatory content bits to check against overlong encoding as opposed to 5 for all longer length code units.</p>\n<h2 id=\"anatomy\">Anatomy</h2>\n<p>Here is anatomical diagram of the example 22-byte code unit from the tldr.</p>\n<p>See the glossary for more information on the definitions of the terms.</p>\n<p>Byte number four is exciting! It is a continuation byte, a start byte, the final start byte, has content bits, and has only some of the mandatory content bits, which are straddled across the final start byte and first non-start byte.</p>\n<p>The main contribution of UTF-8000&#39;s specification is clarity on splitting the highest bits of the first byte of UTF-8 code units into self-synchronization bits and start bits, and then making it clear how to stripe the start bits across the continuation bytes if needed, to achieve arbitrarily large code units.</p>\n<h2 id=\"glossary\">Glossary</h2>\n<p>These terms are ordered somewhat by chronology of first requirement, rather than alphabetically, for convenience.</p>\n<p>Terms used within definitions are underlined clickable hyperlinks.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Term</th><th>Definition</th></tr></thead><tbody><tr><td>Codepoint</td><td>A non-negative integer, aka an unsigned integer.</td></tr><tr><td>Code Unit</td><td>A sequence of UTF-8000 bytes that encode a single codepoint.</td></tr><tr><td>First Byte</td><td>The first, one and only, byte that begins a UTF-8000 code unit.                          The self-synchronization prefix of a first byte is either                         <code>0</code> for ASCII or<code>11</code> for multi-byte code units. This term is <strong>not</strong> synonymous with start byte. A first byte is necessarily a start byte, but not the other way around. It is for this reason that first byte is sometimes also known as first start byte.                          Fun observation: because of the self-synchronization prefix                         <code>0</code> the upper hex nibble of ASCII bytes can only be one of<code>0</code> ,<code>1</code> ,<code>2</code> ,<code>3</code> ,<code>4</code> ,<code>5</code> ,<code>6</code> ,<code>7</code> . This term is mutually exclusive with continuation byte due to self-synchronization.</td></tr><tr><td>Continuation Byte</td><td>A byte beyond the first byte of a multi-byte UTF-8000 code unit.                          The self-synchronization prefix of a continuation byte is <code>10</code> , which is also known as the continuation prefix bits.                          Fun observation: because of the self-synchronization prefix                         <code>10</code> the upper hex nibble of continuation bytes can only be one of<code>8</code> ,<code>9</code> ,<code>A</code> ,<code>B</code> . This term is mutually exclusive with first byte due to self-synchronization.</td></tr><tr><td>Self-Synchronization Prefix</td><td>The highest bits of every UTF-8000 byte that indicate whether it is a first byte or a continuation byte. The possible self-synchronization prefixes form a prefix-free tree: <code>   .----</code> <code>0</code> First byte for ASCII   <code>----</code>1<code> ---</code>0<code>Continuation byte for multi-byte UTF-8000        </code>----<code>1</code> First byte for multi-byte UTF-8000This piece of the clever architecture of UTF-8, which UTF-8000 inherits, provides the property of self-synchronization at a byte level: we can instantaneously tell what kind of byte we are looking at, and where it should belong in a code unit, just by looking at these highest bits.                          This is most useful when decoding part of a file encoded in UTF-8000. If we randomly seek through the file to an arbitrary byte, we can unambiguously tell whether we are at a first byte whence we can begin decoding a new code unit immediately, or that we are at a continuation byte whence we need to seek a little further on in order to find the next first byte in order to begin decoding. Nor do we have to process any bytes prior to our seek position in order to discover some global state or the context of the byte we have seek-ed to; a first byte is always unambiguously a first byte wherever it appears, which we can deduce by its self-synchronization prefix being either <code>0</code> or<code>11</code> . This is useful not only for random access, but also for error recovery. Suppose that we are decoding an error-prone stream of UTF-8000 bytes and that whenever when we encounter an error (e.g. a rogue 0xC0 byte) we wish to <code>U+FFFD</code> See the Wikipedia article for self-synchronizing code for more general info. These bits are highlighted in bright cyan.</td></tr><tr><td>Start Byte</td><td>A byte containing one or more start bits. The start bytes exist contiguously at the beginning of a UTF-8000 code unit. The power of UTF-8000 is that we can have multiple start bytes, to achieve arbitrary code unit lengths, to encode arbitrarily large codepoints. Sometimes it is sensible to colloquially also include ASCII as a start byte when we are talking about the bytes towards the start of a code unit, even though ASCII bytes have no start bits. Every non-ASCII code unit has at least one start byte. The first start byte is the first byte, and it is followed by zero or more continuation bytes that are also start bytes. Therefore because a UTF-8000 code unit can have multiple start bytes, this term is <strong>not</strong> synonymous with first byte.                          In restricting to only UTF-8 without UTF-8000, this term                         <em>is</em> synonymous with first byte. This is because UTF-8-length code units only require one start byte, whether using up to 4 bytes in the current UTF-8 standard (RFC 3629 (2003)), or using up to 6 bytes in former standards (RFC 2044                         (1996) and                         RFC 2279                         (1998)).</td></tr><tr><td>Start Bits</td><td>The unary-code sequence of bits contained in the start bytes of a multi-byte UTF-8000 code unit that tells us the length of the code unit in bytes.                          For a code unit made of <code>n</code> bytes the start bits are<code>n-2</code><code>1</code> bits followed by a terminating<code>0</code> bit. To be clear, the start bits include this terminating zero bit. Thus the start bits sequence is of length<code>n-1</code> and looks like<code>111...10</code> . The possible start bits sequences form a prefix-free tree: <code>   .----</code> <code>0</code> Two byte UTF-8   <code>----</code>1<code> ---</code>0<code>Three byte UTF-8        </code>----<code>1</code> ---<code>0</code> Four byte UTF-8               <code>----</code>1<code> ---</code>0<code>Five byte UTF-8000                    </code>----...    n byte UTF-8000For an <code>n</code> byte code unit where<code>n &lt; 8</code> the start bits all fit together snugly in the first byte. Otherwise they are striped across as many of the first few bytes as they need, filling the free bits that are not occupied by continuation prefix bits.                          This is another piece of the clever architecture of UTF-8, which UTF-8000 inherits, that provides the property of self-punctuation also known as a <code>0</code> bit, we know exactly how many bytes we expect in that code unit. Notwithstanding errors we can therefore succeed in decoding the code unit by reading exactly that many bytes, and no more. This avoids a problem of dumber variable-length encodings whose code units do not intrinsically indicate their length: one has to read beyond the last byte of a code unit, that is one reads the first byte of the next code unit, in order to know that the current code unit has finished. For very dumb encodings which have neither self-synchronization nor self-punctuation, to make random access possible one would have to put dedicated auxiliary bytes,  See the Wikipedia articles for prefix code and unary coding for more general info. This term is mutually exclusive with content bits. These bits are highlighted in bright magenta.</td></tr><tr><td>Content Byte</td><td>A byte containing one or more content bits.                          A byte being a content byte does not imply that it is a continuation byte. For example a 3-byte code unit begins with <code>1110xxxx</code> , which contains 4 content bits and is not a continuation byte.                          A byte being a continuation byte does not imply that it is a content byte. For example a 22-byte code unit contains                         <code>10111111</code> as its second byte, which is a continuation byte and has no content bits.</td></tr><tr><td>Content Bits</td><td>The sequence of bits in a code unit beyond the start bits and to the end of the code unit, in which the codepoint&#39;s binary bits are stored. For example a 3-byte code unit, which has the form                         <code>1110xxxx</code><code>10xxxxxx</code><code>10xxxxxx</code> , has 16 content bits.                          For ASCII there are 7 content bits. These seven bits <code>xxxxxxx</code> combined with a byte&#39;s highest bit being set to the self-synchronization prefix<code>0</code> means that ASCII is perfectly included into UTF-8 without being altered. Thus ASCII code units take the form<code>0xxxxxxx</code> . Otherwise for an <code>n</code> byte code unit, where<code>n &gt; 1</code> , there are<code>5n+1</code> content bits. This is how we arrive at that formula: We start with<code>n</code> blank bytes, each of which has<code>8</code> bits. For each byte<code>2</code> bits are taken by the self-synchronization prefix. Then an additional<code>n-1</code> bits are taken by the start bits. Thus there are<code>8n - 2n - (n-1) = 5n+1</code> bits left for content bits. Another way to think about the<code>5</code> in this formula is by extending from<code>n-1</code> bytes to<code>n</code> bytes by appending another continuation byte. By doing this we gain<code>6</code> free bits in the continuation byte, but we lose<code>1</code> bit to the longer start bits sequence, thus overall we gain<code>6-1 = 5</code> bits for content bits. This term is mutually exclusive with start bits. These bits are highlighted in lime.</td></tr><tr><td>Mandatory Content Byte</td><td>A byte containing one or more mandatory content bits. These are the bytes we check for overlong encoding when decoding a code unit.</td></tr><tr><td>Mandatory Content Bits</td><td>The first 0, 4, or 5 content bits of a code unit in which there must be at least one <code>1</code> bit, lest the bytes form an overlong encoding, which is forbidden. For ASCII there are 0 mandatory content bits, and thus no anti-overlong checking is required. This is because ASCII is the smallest possible code unit. For 2-byte UTF-8000 there are 4 mandatory content bits. This is because in the jump from 1-byte ASCII to 2-byte UTF-8 we jump from 7 content bits to 11 content bits. Thus the number of content bits we gain is 11 minus 7 which is 4. Otherwise for <code>n</code> byte UTF-8000, where<code>n &gt; 2</code> , there are 5 mandatory content bits. This is because in the jump from<code>n-1</code> byte UTF-8000 to<code>n</code> byte UTF-8000 we add on an extra continuation byte, which has 6 free bits, but we lose 1 bit to the longer start bits sequence. Thus overall the number of content bits we gain is 6 minus 1 which is 5. Read about overlong encoding for why mandatory content bits are of interest. These bits are highlighted in bright lime.</td></tr><tr><td>Overlong Encoding</td><td>Forbidden                           For example one could incorrectly try to encode the codepoint 0x41, 65, ASCII capital A, using 2-byte UTF-8 as                         <code>11000001</code><code>10000001</code> . Observe that all the mandatory content bits are<code>0</code> which is the definition an overlong encoding. This indicates that we could have encoded 0x41 in a shorter code unit, in this case as ASCII<code>01000001</code> .                          Security is one main reason why we forbid overlong encoding. For example we ensure that                         <code>11100000</code><code>10000000</code><code>10000000</code> cannot be decoded as codepoint 0, the null<code>strcpy(3)</code> and friends.<code>strcpy</code> would not interpret this Uniqueness of encoding is another reason why we forbid overlong encoding. Every codepoint has one unique valid representation as a UTF-8000 code unit, which is easy to encode and decode using bitshifting.                          Fun observation: because all 4 of 2-byte UTF-8&#39;s mandatory content bits lie in the first-and-final start byte, we can explicitly rule out                         <code>11000000</code> (0xC0) and<code>11000001</code> (0xC1) as permanently invalid bytes. They will never ever appear anywhere in a valid UTF-8000                         code unit!</td></tr></tbody></table></div>\n<h2 id=\"properties\">Properties</h2>\n<p>Many of these properties of UTF-8000 are explained in detail in an appropriate section of the glossary and hyperlinks to the glossary are provided.</p>\n<h3 id=\"bit-counts\">Bit Counts</h3>\n<p>The number of content bits and mandatory content bits are very predictable as a function of <code>n</code>, the length of a code unit.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>code unit length</th><th>number of content bits</th><th>number of mandatory content bits</th></tr></thead><tbody><tr><td><code>n = 1</code></td><td><code>7</code></td><td><code>0</code></td></tr><tr><td><code>n = 2</code></td><td><code>5n+1 ( = 11)</code></td><td><code>4</code></td></tr><tr><td><code>n &gt; 2</code></td><td><code>5n+1</code></td><td><code>5</code></td></tr></tbody></table></div>\n<h3 id=\"why-the-special-cases\">Why the Special Cases?</h3>\n<p>As stated in the tldr, there are only two special cases, both of which are inherited from UTF-8:</p>\n<p><strong>1-byte UTF-8 (ASCII)</strong> which has two points of interest:</p>\n<ul><li>It has 7 content bits which does not fit the pattern of <code>5n+1</code> . See the glossary section for content bits for an explanation, and see the rejected alternative ASCVI code for a version of UTF-8 if ASCII were 6 bit instead of 7 bit which eliminates this special case.</li><li>ASCII has 0 mandatory content bits because it cannot possibly be overlong since it is the smallest possible code unit. This is fine.</li></ul>\n<p><strong>2-byte UTF-8</strong> which has one point of interest:</p>\n<ul><li>It has 4 mandatory content bits, as opposed to 5 for all longer code units. See the glossary section for mandatory content bits for an explanation.</li></ul>\n<p>The remarkable fact that UTF-8000 does not introduce any new special cases in extending UTF-8 is confirmation to me that this is the canonical, correct way to extend UTF-8. In other words UTF-8 in its current restricted 4 byte form <em>is</em> UTF-8000, but only a small part of it.</p>\n<pre><code>            The fact that we are even able to extend in the first place is also testament to the clever planning and care that Ken Thompson and Rob Pike put into the architecture of UTF-8, which we ensure to maintain as we extend to UTF-8000. Unary code codewords for the start bits sequences, which form a self-similar tree, were a great choice being simple and extensible. In the earliest draft of UTF-8, the six-byte start-byte looked like `111111xx`. This was changed a few days later to `1111110x`. That way the number of content bits is not a special case, and the start bits don&#39;t saturate the unary code binary tree, leaving the door open for our future expansion.</code></pre>\n<p>This is why I think of UTF-8 as the capstone of the Unix Philosophy.</p>\n<h3 id=\"information-rate\">Information Rate</h3>\n<p>What proportion of a code unit is content bits?</p>\n<p>For ASCII this is <code>7/8 = 87.5%</code>.</p>\n<p>Otherwise for an <code>n</code> byte code unit this is <code>(5n+1) / 8n</code>, that is <code>5n+1</code> content bits out of a total of <code>8n</code> bits from <code>n</code> bytes. We can rewrite this as <code>(5/8) + 1/(8n)</code> which moderately quickly approaches <code>5/8 = 62.5%</code>. It is nice that this limit is nonzero and does not depend on <code>n</code>.</p>\n<h3 id=\"self-synchronization\">Self-Synchronization</h3>\n<p>Inherited from UTF-8 and maintained in UTF-8000.</p>\n<p>See the glossary section for self-synchronization prefix for an explanation of self-synchronization.</p>\n<pre><code>            Here&#39;s a bit of history: Self-synchronization is one of the reasons why Ken Thompson and Rob Pike decided to design UTF-8, to supersede the earlier FSS-UTF draft by Dave Prosser et al. FSS-UTF proposed a design like eg\n            `110xxxxx`\n            `1xxxxxxx`\n            `1xxxxxxx`\n            for three-byte code units. The problem with it is that one cannot distinguish between first bytes (`110xxxxx`) and continuation bytes (`110xxxxx`) without knowing the prior history of a stream. The UTF-8 fix is to make first byte and continuation byte values disjoint from each other, as one can witness in the byte map below. I have not put Prosser&#39;s draft into the rejected ideas section as it has already been formally addressed and superseded by UTF-8.</code></pre>\n<h3 id=\"self-punctuation\">Self-Punctuation</h3>\n<p>Inherited from UTF-8 and maintained in UTF-8000.</p>\n<p>See the glossary section for start bits for an explanation of self-punctuation.</p>\n<h3 id=\"byte-map\">Byte Map</h3>\n<p>Extended from UTF-8, making use of the higher value bytes. Based off Wikipedia&#39;s UTF-8 Byte Map.</p>\n<div class=\"table-wrap\"><table><thead><tr><th></th><th>0</th><th>1</th><th>2</th><th>3</th><th>4</th><th>5</th><th>6</th><th>7</th><th>8</th><th>9</th><th>A</th><th>B</th><th>C</th><th>D</th><th>E</th><th>F</th></tr></thead><tbody><tr><td>0</td><td>␀</td><td>␁</td><td>␂</td><td>␃</td><td>␄</td><td>␅</td><td>␆</td><td>␇</td><td>␈</td><td>␉</td><td>␊</td><td>␋</td><td>␌</td><td>␍</td><td>␎</td><td>␏</td></tr><tr><td>1</td><td>␐</td><td>␑</td><td>␒</td><td>␓</td><td>␔</td><td>␕</td><td>␖</td><td>␗</td><td>␘</td><td>␙</td><td>␚</td><td>␛</td><td>␜</td><td>␝</td><td>␞</td><td>␟</td></tr><tr><td>2</td><td>␠</td><td>!</td><td>&quot;</td><td>#</td><td>$</td><td>%</td><td>&amp;</td><td>&#39;</td><td>(</td><td>)</td><td>*</td><td>+</td><td>,</td><td>-</td><td>.</td><td>/</td></tr><tr><td>3</td><td>0</td><td>1</td><td>2</td><td>3</td><td>4</td><td>5</td><td>6</td><td>7</td><td>8</td><td>9</td><td>:</td><td>;</td><td>&lt;</td><td>=</td><td>&gt;</td><td>?</td></tr><tr><td>4</td><td>@</td><td>A</td><td>B</td><td>C</td><td>D</td><td>E</td><td>F</td><td>G</td><td>H</td><td>I</td><td>J</td><td>K</td><td>L</td><td>M</td><td>N</td><td>O</td></tr><tr><td>5</td><td>P</td><td>Q</td><td>R</td><td>S</td><td>T</td><td>U</td><td>V</td><td>W</td><td>X</td><td>Y</td><td>Z</td><td>[</td><td>\\</td><td>]</td><td>^</td><td>_</td></tr><tr><td>6</td><td>`</td><td>a</td><td>b</td><td>c</td><td>d</td><td>e</td><td>f</td><td>g</td><td>h</td><td>i</td><td>j</td><td>k</td><td>l</td><td>m</td><td>n</td><td>o</td></tr><tr><td>7</td><td>p</td><td>q</td><td>r</td><td>s</td><td>t</td><td>u</td><td>v</td><td>w</td><td>x</td><td>y</td><td>z</td><td>{</td><td>|</td><td>}</td><td>~</td><td>␡</td></tr><tr><td>8</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>9</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>A</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>B</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>C</td><td></td><td></td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td></tr><tr><td>D</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td></tr><tr><td>E</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td></tr><tr><td>F</td><td>4</td><td>4</td><td>4</td><td>4</td><td>4</td><td>4</td><td>4</td><td>4</td><td>5</td><td>5</td><td>5</td><td>5</td><td>6</td><td>6</td><td>7</td><td>8+</td></tr></tbody></table></div>\n<p>All bytes except 0xC0 and 0xC1, colored in tomato red, can appear in a valid UTF-8000 stream. See the glossary section for overlong encoding for an explanation of why those two bytes never appear.</p>\n<p>ASCII, colored in gold yellow, occupies the first half of the table, being 7-bit. Continuation bytes occupy the region colored in sandybrown orange. All other bytes are first bytes for multi-byte code units, whose lengths are indicated in the table.</p>\n<h3 id=\"strcmp-3-ordering\"><code>strcmp(3)</code> Ordering</h3>\n<pre><code>          Inherited from UTF-8 and maintained in UTF-8000.\n\n            The self-synchronization prefixes of first bytes are monotonically-increasing-ly ordered with respect to code unit length. In other words ASCII is of length 1 and multi-byte is of length greater than 1, and `0` &lt; `11` occupying the highest bits of UTF-8000 bytes.\n          \n\n            The start bit sequences are also monotonically-increasing-ly ordered with respect to code unit length. In other words if `n &lt; m` then `111...[n]...10` `&lt;` `111...[m]...10` as an integer value, occupying the heads of the code unit bytes beyond the self-synchronization prefixes. This would not have been the case had UTF-8 been designed to use the alternative form of unary codewords given by `000...01`.</code></pre>\n<p>The content bits of code units are also monotonically-increasing-ly ordered with respect to codepoint value.</p>\n<p>Combining these three things together means that <code>strcmp(3)</code>, the C stdlib string comparing function, works the same way on UTF-8000 bytes as it does on UTF-8, as it does on ASCII, effectively comparing the encoded codepoint values against each other without having to actually decode the code units. Nice!</p>\n<h3 id=\"no-endianness\">No Endianness</h3>\n<p>The quantum</p>\n<p> of ASCII, UTF-8, and UTF-8000 is a single byte. This makes life a breeze! There is no need for a concept of endianness for UTF-8000.</p>\n<p>UTF-16 however has a quantum of two bytes, 16-bit units. When writing the codewords in bytes, 8-bit units, should the byte containing the most significant digits or least significant digits be written first? Big-endian, or little-endian? This choice gives UTF-16 two variants, UTF-16-BE and UTF-16-LE. If one cannot predetermine the endianness of a stream, one may wish to use a BOM which is discussed below.</p>\n<h3 id=\"bom-support\">BOM Support</h3>\n<p>A Byte Order Mark (BOM) is used at the start of an encoded text stream to indicate what encoding is used. I have never actively used BOMs myself so I&#39;ve only put a bit of thought into this section.</p>\n<p>As far as I&#39;m aware we don&#39;t break BOM support for UTF-8, though we may wish to have a different BOM to strictly distinguish UTF-8 from UTF-8000. Maybe UTF-8000 could have multiple BOMs, one for each integer N greater than or equal to four, to indicate to a decoder the maximum expected code unit length.</p>\n<pre><code>            One of the reasons why `U+FFFE` is not a valid Unicode Scalar Value is because `0xFE 0xFF` is the BOM for UTF-16. Since UTF-16 code units are two bytes wide, one may read either `0xFE 0xFF` or `0xFF 0xFE` depending on endianness. To make it clear that `0xFF 0xFE` implies correct for endianness</code></pre>\n<p> and cannot be mistaken for a legitimate character, <code>U+FFFE</code> is designated as <code>&lt;noncharacter-FFFE&gt;</code>. We are relieved in that neither <code>11111111</code> <code>11111110</code> nor <code>11111110</code> <code>11111111</code> are valid UTF-8000 sequence extracts, ie UTF-8000 does not introduce incompatibilities with UTF-16.</p>\n<h3 id=\"arbitrary-lengths-sensible-limits\">Arbitrary Lengths, Sensible Limits</h3>\n<p>I think we&#39;ve made it clear by now that UTF-8000 code units can be arbitrarily large. In practice however one <em>may</em> wish to set a sensible limit on code unit lengths when decoding. Here we&#39;ll discuss a method of finding some nice code unit lengths whose code units store <code>5n+1 = 2^N</code> bits, as we are often interested in powers of 2 in computer science.</p>\n<p>It is a common observation that 3-byte UTF-8 stores <code>5 * 3 + 1 = 16</code> bits, meaning the Basic Multilingual Plane of Unicode can be encoded in one two and three byte UTF-8. We see that <code>2^4 mod5 = 16 mod5 = 1 mod5</code>; if <code>5n+1</code> is to be <code>2^N</code> for some <code>n</code> then certainly <code>2^N = 1 mod5</code>. If we enumerate powers of two modulo five then there is a very predictable repeating pattern of 1, 2, 4, 3</p>\n<p>. Formally you might say that 2 is a generator of 𝔽&lt;sub&gt;5&lt;/sub&gt;&lt;sup&gt;*&lt;/sup&gt; if you want impress a mathematician! The takeaway is that when <code>N=4K</code> for <code>K≥1</code> we can find a corresponding <code>n</code> such that <code>5n+1 = 2^N</code>. We can rewrite <code>2^N</code> as <code>2^(4K) = 16^K</code>.</p>\n<p>In other words any power of 16 has a UTF-8000 code unit length containing that many bits</p>\n<p>. Here are a few of these for reference.</p>\n<div class=\"table-wrap\"><table><thead><tr><th><code>K</code></th><th><code>N=4K</code></th><th>number of content bits <code>= 2^N</code></th><th>code unit length <code>= (2^N - 1) / 5</code></th></tr></thead><tbody><tr><td><code>1</code></td><td><code>4</code></td><td><code>16</code></td><td><code>3</code></td></tr><tr><td><code>2</code></td><td><code>8</code></td><td><code>256</code></td><td><code>51</code></td></tr><tr><td><code>3</code></td><td><code>12</code></td><td><code>4096</code></td><td><code>819</code></td></tr><tr><td><code>4</code></td><td><code>16</code></td><td><code>65536</code></td><td><code>13107</code></td></tr><tr><td><code>...</code></td><td><code>...</code></td><td></td><td></td></tr></tbody></table></div>\n<p>Do remember that strictly speaking one shouldn&#39;t allow overlong encodings, if one were for example thinking of storing a small <code>uint256_t</code> key in a 51 byte code unit! UTF-8000&#39;s variable width nature helps out leading to smaller code units for smaller integers.</p>\n<h2 id=\"intuitive-derivation\">Intuitive Derivation</h2>\n<p>There are a few ways that one could arrive at the design for UTF-8000 and the bit counts above.</p>\n<p>One may think to start with UTF-8, notice that the start byte of an <code>n</code> byte code unit is prefixed with the unary codeword of length <code>n+1</code>, that is <code>n</code> <code>1</code> bits followed by a <code>0</code>, and then figure out how to extend those bits and roll them over into the continuation bytes without losing any important properties. This is what I <em>originally</em> did.</p>\n<p>Writing this document over a couple of weeks made me introspect the code unit anatomy further, whence I figured out that separating the leading bits into a self-synchronization part and self-punctuation part further illuminates and simplifies the thought process. We shall thus proceed with this perspective.</p>\n<h4 id=\"blank-slate\">Blank Slate</h4>\n<p>We set out to derive the design of an <code>n</code> byte code unit, starting out with <code>n</code> blank bytes, all of whose bits could possibly be content bits.</p>\n<pre><code>            `00000000`\n            `00000000`\n            `00000000`\n            `...`\n            `00000000`\n          \n\n            To achieve self-synchronization we need to distinguish the first byte of the code unit from the continuation bytes that follow. We could do that by setting the highest bit of first bytes to a\n            `0` and to `1` for continuation bytes. Doing it this way round maintains compatibility with ASCII&#39;s highest bit being `0`.\n          \n\n            `00000000`\n            `10000000`\n            `10000000`\n            `...`\n            `10000000`</code></pre>\n<p>With the design so far, all code units begin with an ASCII byte. When decoding a code unit, we have no idea whether this first byte actually is ASCII, or it is the first byte of a multi-byte code unit. We want self-punctuation, where a code unit intrinsically tells us how long it is.</p>\n<p>To achieve self-punctuation we create a prefix-free code binary tree, whose leaf node codewords correspond to code unit lengths. These are the start bits sequences. The codeword for <code>n</code> shall be embedded inside the code unit towards the start. It must therefore be short enough to fit into the <code>n</code> bytes, and reasonably computationally predictable. We try:</p>\n<pre><code>  .----</code></pre>\n<p><code>0</code>                          One byte UTF-8 (ASCII)\n  <code>----</code>1<code>---</code>0<code>                    Two byte UTF-8\n        </code>----<code>1</code>---<code>0</code>            Three byte UTF-8\n              <code>----</code>1<code>---</code>0<code>       Four byte UTF-8\n                    </code>----<code>1</code>---<code>0</code> Five byte UTF-8000\n                          `----...    n byte UTF-8000\n              This seems reasonably simple so far. We stripe the start bits into the available bits not taken by the self-synchronization prefix. All other bits shall be content bits.</p>\n<p>| 1 | <code>00xxxxxx</code> |  |  |  |  |  | \n| 2 | <code>010xxxxx</code> | <code>1xxxxxxx</code> |  |  |  |  | \n| 3 | <code>0110xxxx</code> | <code>1xxxxxxx</code> | <code>1xxxxxxx</code> |  |  |  | \n| ... |  |  |  |  |  |  | \n| 17 | <code>01111111</code> | <code>11111111</code> | <code>1110xxxx</code> | <code>1xxxxxxx</code> | ... | <code>1xxxxxxx</code> | \n| ... |  |  |  |  |  |  | </p>\n<pre><code>            But wait we&#39;ve broken the distinction of ASCII! We cannot tell the difference between eg\n            `0110xxxx`\n            and\n            `0110xxxx`, or\n            `01111111`\n            and\n            `01111111`. This code would only work if ASCII were six-bit instead of seven-bit. Out of curiosity we investigate this code in the rejected alternatives section ASCVI</code></pre>\n<p>.</p>\n<pre><code>            To maintain compatibility with ASCII we must treat it as a special case, whereby the self-synchronization prefix `0` is alone sufficient to characterize ASCII. This highlights that the architecting of UTF-8 was not purely a mathematics problem, but was also an engineering problem, working around what already exists.\n          \n\n            Seeing the ASCII-characterizing prefix `0` and the erstwhile continuation prefix `1` as forming a prefix-free tree, albeit only of size two, we must repurpose the the latter codeword as the beginning of the self-synchronization prefixes for first bytes and continuation bytes of multi-byte code units. We choose our new self-synchronization prefixes as `11` for start bytes and `10` for continuation bytes. This produces the following tree:</code></pre>\n<pre><code>  .----</code></pre>\n<p><code>0</code>              First byte for ASCII\n  <code>----</code>1<code>---</code>0<code> Continuation byte for multi-byte UTF-8000\n        </code>----<code>1</code>        First byte for multi-byte UTF-8000\n              Accordingly adjusting the self-punctuation codewords to apply only to multi-byte code units produces the following tree:</p>\n<pre><code>  .----</code></pre>\n<p><code>0</code>                    Two byte UTF-8\n  <code>----</code>1<code>---</code>0<code>            Three byte UTF-8\n        </code>----<code>1</code>---<code>0</code>       Four byte UTF-8\n              <code>----</code>1<code>---</code>0<code> Five byte UTF-8000\n                    </code>----...    n byte UTF-8000\n              Putting these mechanisms together yields UTF-8000 and we&#39;re done!</p>\n<p>| ASCII |  |  |  |  |  |  |  |  |  |  |  | \n| 1 | <code>0xxxxxxx</code> |  |  |  |  |  |  |  |  |  |  | \n| UTF-8 |  |  |  |  |  |  |  |  |  |  |  | \n| 2 | <code>110xxxxx</code> | <code>10xxxxxx</code> |  |  |  |  |  |  |  |  |  | \n| 3 | <code>1110xxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  |  |  |  |  |  |  | \n| 4 | <code>11110xxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  |  |  |  |  |  | \n| UTF-8000 |  |  |  |  |  |  |  |  |  |  |  | \n| 5 | <code>111110xx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  |  |  |  |  | \n| 6 | <code>1111110x</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  |  |  |  | \n| 7 | <code>11111110</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  |  |  | \n| 8 | <code>11111111</code> | <code>100xxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  |  | \n| 9 | <code>11111111</code> | <code>1010xxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  | \n| 10 | <code>11111111</code> | <code>10110xxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  | \n| ... |  |  |  |  |  |  |  |  |  |  |  | \n| 22 | <code>11111111</code> | <code>10111111</code> | <code>10111111</code> | <code>10110xxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | ... | <code>10xxxxxx</code> | \n| ... |  |  |  |  |  |  |  |  |  |  |  | </p>\n<p>It is a trivial* result of coding theory that the product of two prefix-free codes is also a prefix-free code. The product of our trees looks like:</p>\n<pre><code>  .----</code></pre>\n<p><code>0</code>                                                  ASCII byte\n  <code>----</code>1<code>---</code>0<code>                               UTF-8 continuation byte\n        </code>----<code>1</code>---<code>0</code>                    Two byte UTF-8    start byte\n              <code>----</code>1<code>---</code>0<code>            Three byte UTF-8    start byte\n                    </code>----<code>1</code>---<code>0</code>       Four byte UTF-8    start byte\n                          <code>----</code>1<code>---</code>0<code> Five byte UTF-8000 start byte\n                                </code>----...    n byte UTF-8000 start byte\n              Without the color highlighting this is how most people think about UTF-8: a start byte whose prefix of <code>n</code> <code>1</code> bits and a terminating <code>0</code> bit provides both self-synchronization and self-punctuation, and continuation bytes using a short prefix of <code>10</code> for space-efficient encoding. This makes sense for a small number of bytes, but the trick to unlock a perspective of infinite extensibility is to split this tree into the self-synchronization part and self-punctuation part; ie we un-product those prefix-free codes. Failing to do this leads to the rejected alternative UTF-Infinity.</p>\n<h2 id=\"encoding\">Encoding</h2>\n<p>This section is based off the reference implementation which is written in Python. It is well documented, and is more specific on how to use bitwise operations. This is an abridged HTML version.</p>\n<p>Suppose that we have an unsigned integer <code>n</code> that we want to encode in UTF-8000. Initialize an empty dynamic array of bytes <code>ret_ints</code> that will store the UTF-8000 code unit.</p>\n<p>If <code>n &lt; 0x80</code>, eg <code>n = 0x41</code>, then insert <code>n</code> at the head of <code>ret_ints</code> and we are done. This is the ASCII byte for <code>n</code>, which in our example of <code>n = 0x41</code> is a capital letter a, <code>&#39;A&#39;</code>.</p>\n<p>Otherwise for <code>n ≥ 0x80</code>, eg <code>n = 0x0321C0FFEE8086</code>, we use UTF-8000. Initialize an integer counter <code>n_bits_content_occupied</code> to zero.</p>\n<pre><code>            Our example `n`&#39;s content bits look like `11 001000 011100 000011 111111 111011 101000 000010 000110` as a big raw number, with spaces added for visual ease.\n          \n\n            While `n` has more than 6 content bits, aka `n &gt; 63 =` `00111111`, extract the least-significant 6 bits of `n` and insert them at the head of `ret_ints`, incrementing `n_bits_content_occupied` by 6 and downwards bitshifting `n` by 6.\n          \n\n            `n_bits_content_occupied = 48`, `n = 0b``11`,</code></pre>\n<p><code>ret_ints</code>:\n                <code>00001000</code>\n                <code>00011100</code>\n                <code>00000011</code>\n                <code>00111111</code>\n                <code>00111011</code>\n                <code>00101000</code>\n                <code>00000010</code>\n                <code>00000110</code></p>\n<pre><code>          Now insert the rest of `n` at the head of `ret_ints`. Count the number of bits left in `n` by downwards bitshifting `n` one bit at a time while it is non-zero. This is between 1 and 6 (inclusive), which we also add to `n_bits_content_occupied`.</code></pre>\n<p><code>n_bits_content_occupied = 50</code>, <code>n = 0</code>,</p>\n<p><code>ret_ints</code>:\n                <code>00000011</code>\n                <code>00001000</code>\n                <code>00011100</code>\n                <code>00000011</code>\n                <code>00111111</code>\n                <code>00111011</code>\n                <code>00101000</code>\n                <code>00000010</code>\n                <code>00000110</code></p>\n<pre><code>          Now we calculate how many bytes our UTF-8000 code unit requires, `n_utf_8000_bytes_needed`. We know that a `k`-byte code unit has capacity for `5k+1` content bits. Therefore `⌈(n_bits_content_occupied-1) / 5⌉` is the sufficient and minimal answer. Any larger code unit size would lead to an overlong encoding! For our example `n_utf_8000_bytes_needed = ⌈(50-1) / 5⌉ = 10`.</code></pre>\n<p>Leftwards pad <code>ret_ints</code> with empty bytes to the length <code>n_utf_8000_bytes_needed</code>.</p>\n<p><code>ret_ints</code>:\n                <code>00000000</code>\n                <code>00000011</code>\n                <code>00001000</code>\n                <code>00011100</code>\n                <code>00000011</code>\n                <code>00111111</code>\n                <code>00111011</code>\n                <code>00101000</code>\n                <code>00000010</code>\n                <code>00000110</code></p>\n<pre><code>            Now we add the start bits. The number of `1` start bits is equal to two less than the number of bytes in the code unit, which we just calculated. We therefore calculate `q, r = divmod(n_utf_8000_bytes_needed-2, 6)`, which tells us we need `q` hextets full of `1` start bits, and a final hextet of zero to five `1` bits, which also has space to contain the terminating `0` bit. In our example `(q = 1, r = 2) = divmod(10-2, 6)`.</code></pre>\n<p>Apply the start bits across <code>ret_ints</code> using bitwise-or. The final start bits hextet can be given by <code>((1 &lt;&lt; r) - 1) &lt;&lt; (6 - r)</code>.</p>\n<p><code>ret_ints</code>:\n                <code>00111111</code>\n                <code>00110011</code>\n                <code>00001000</code>\n                <code>00011100</code>\n                <code>00000011</code>\n                <code>00111111</code>\n                <code>00111011</code>\n                <code>00101000</code>\n                <code>00000010</code>\n                <code>00000110</code></p>\n<pre><code>          Any of the lowest six bits of each byte that are not set by this point, unoccupied by content bits and untouched by start bits, are really content bits that the `k`-byte capacity provides but that we didn&#39;t need. Our example&#39;s `n_bits_content_occupied = 50` is one less than `5k+1 = 5*10+1 = 51`. We can color highlight it green as a content bit for completion&#39;s sake.</code></pre>\n<p><code>ret_ints</code>:\n                <code>00111111</code>\n                <code>00110011</code>\n                <code>00001000</code>\n                <code>00011100</code>\n                <code>00000011</code>\n                <code>00111111</code>\n                <code>00111011</code>\n                <code>00101000</code>\n                <code>00000010</code>\n                <code>00000110</code></p>\n<pre><code>            Now we crown the bytes with their self-synchronization prefixes, which delivers us from hextets to UTF-8000 octets. The first byte&#39;s self-synchronization prefix is\n            `11`, and continuation bytes have `10`.</code></pre>\n<p><code>ret_ints</code>:\n                <code>11111111</code>\n                <code>10110011</code>\n                <code>10001000</code>\n                <code>10011100</code>\n                <code>10000011</code>\n                <code>10111111</code>\n                <code>10111011</code>\n                <code>10101000</code>\n                <code>10000010</code>\n                <code>10000110</code></p>\n<pre><code>          And we&#39;re done!</code></pre>\n<h2 id=\"decoding\">Decoding</h2>\n<p>As with the encoding section, this section is based off the reference implementation which is written in Python and well documented.</p>\n<p>Suppose that we are receiving a stream of UTF-8000 bytes (possibly with errors!), and we wish to extract and taxonomically annotate the incoming code units. There are a few ways that we could approach this, such as using the byte map as a state machine, which I want to try in the future, or the classic way of using bitwise masks. We are going to use the latter approach in this section, as we describe how to decode a single code unit. But first, a look at error handling.</p>\n<h3 id=\"error-recovery\">Error Recovery</h3>\n<p>The errors that can occur when decoding a UTF-8000 stream are:</p>\n<ol><li><p></p><pre><code>            Reading a continuation byte (`10` ) when we are expecting the first byte of a code unit (`0` or`11` ).</code></pre></li><li><p></p><pre><code>            Reading a first byte (`0` or`11` ) when we are expecting a continuation byte (`10` ).</code></pre></li><li>Early EOF midway through a code unit.</li><li><p></p><pre><code>            Encountering an overlong encoding</code></pre>\n<ol><li><p></p><pre><code>              For 2-byte code units this is bytes 0xC0 (`11000000` ) and 0xC1 (`11000001` ).</code></pre></li><li><p></p><pre><code>              For `n` -byte code units in general, with`n &gt; 2` , eg`11100000``10010111``10010000` .</code></pre></li></ol></li><li><p></p><pre><code>                For 2-byte code units this is bytes 0xC0 (</code></pre></li><li>Encountering an encoded surrogate codepoint in the range <code>U+D800</code> to<code>U+DFFF</code> , which is forbidden for compatibility with UTF-16.</li></ol>\n<p>For standard UTF-8 one would also have to be concerned with codepoints beyond the range <code>U+10FFFF</code> whence bytes 0xF5 to 0xFF go unused.</p>\n<p>For any of these errors a parser could raise an exception and refuse to continue. Alternatively it could take advantage of UTF-8000&#39;s self-synchronization property, and keep calm and carry on</p>\n<p>, yielding Unicode replacement characters <code>U+FFFD</code> �</p>\n<p> until we reach the first byte of the next code unit. Let us investigate the latter course.</p>\n<p>To handle error 1. the parser should return one � and get ready to parse the next code unit. When handling error 2. the parser should make sure to unpop</p>\n<p> the byte encountered, as it is the first byte of the next code unit. When handling errors 2. through to 5. there are a couple of mainstream approaches for yielding � characters:</p>\n<h4 id=\"maximal-subpart\">Maximal Subpart</h4>\n<pre><code>            The Unicode Consortium recommends, but does not enforce, a maximal subpart</code></pre>\n<p> approach, in which the longest well-formed part of a code unit should return a single � character, rather than one for each byte involved. For example the three bytes in error 4.2. above should return one � as it is well-formed with respect to self-synchronization and self-punctuation, and only invalid at an overlong level, being an overlong encoding of\n                <code>11010111</code>\n                <code>10010000</code>\n                <code>U+05D0</code>, a Hebrew letter Aleph &#39;א&#39;.</p>\n<pre><code>            I dislike this approach. Waiting for maximal subparts has the problem that the rest of an invalid code unit may never arrive. If we receive the bytes\n            `11100000`\n            `10010111`\n            from a socket, then the remote end may be waiting for us to chastise their overlong opening bytes with a response, because we can already tell that these bytes form part of an invalid code unit. Using the maximal subpart approach we *also* would be waiting, for the remote end to send a continuation byte eg `10010000` to form an overlong but otherwise complete 3-byte code unit. This is uncooperative, and not what I want.</code></pre>\n<h4 id=\"one-for-each-byte-read\">One � For Each Byte Read</h4>\n<p>We are going to do what Python, my terminal KDE Konsole, and others do, and simply return a � character for each invalid byte. In Python <code>b&#39;\\xE0\\x97\\x90&#39;.decode(errors=&#39;replace&#39;)</code> returns <code>&#39;���&#39;</code>.</p>\n<p>This approach is easier and more versatile. The end user can see how many invalid bytes occurred by counting the number of � characters. There are no deadlock waiting events that can occur with the maximal subpart approach.</p>\n<h3 id=\"the-main-decode-loop\">The Main Decode Loop</h3>\n<p>Initialize an empty dynamic array of bytes <code>parsed_bytes</code> that will store the bytes as we parse them.</p>\n<p>Read a byte, store it as <code>start_byte</code>. Use bitwise masks to find the index, <code>idx_0</code>, of the most-significant zero bit in the byte. If there are no zeros in this byte, <code>idx_0</code> should be set to <code>-1</code>.</p>\n<pre><code>            If `idx_0 == 7` (`0xxxxxxx`) then `start_byte` is an ASCII byte, which has seven content bits. Append `start_byte` to `parsed_bytes` and we are done.\n          \n\n            If `idx_0 == 6` (`10xxxxxx`) then `start_byte` is a continuation byte, which is an invalid start byte. Go to error 1.\n          \n\n            If `idx_0 == 5` (`110xxxxx`) then this is the first byte of a 2-byte code unit. We treat this as a special case because there are only 4 mandatory content bits, not 5. As they are all contained in `start_byte` we can check them immediately for overlong encoding, to see if we need to handle error 4.1. If `start_byte` passes this check then append it to `parsed_bytes` and await a continuation byte. Handle error 2 if necessary, else append the continuation byte to `parsed_bytes` and we&#39;re done.</code></pre>\n<p>We could (should really) make <code>idx_0 == 4</code> a special case too, to check for and forbid the surrogate ranges. I have omitted this for the time being and we drop through to the generic case below.</p>\n<pre><code>            Otherwise we enter the generic case (`111[1...]`). Initialize an integer counter `n_bytes_expected` to 2. Increment `n_bytes_expected` by `5 - idx_0`, as `idx_0` now serves the purpose being the index of the terminating zero of the start bits, `0`.</code></pre>\n<p>If <code>idx_0 == -1</code> then our code unit has multiple start bytes, exciting! Append <code>start_byte</code> to <code>parsed_bytes</code>, and <code>while(1)</code>:</p>\n<pre><code>            Read a byte, and make sure it is a continuation byte lest we go to error 2. Use bitwise masks to find `idx_0`, the index of the most-significant zero bit in the lowest *six* bits of the byte, setting `idx_0` to `-1` if there is none. This is to continue trying to find the `0` start bit. Increment `n_bytes_expected` by `5 - idx_0`. If `idx_0 == -1` then append `start_byte` to `parsed_bytes` and continue again through this loop, until we find the `0` bit, at which point we break this loop.</code></pre>\n<p>At this stage, whether our code unit has multiple start bytes or just one, <code>start_byte</code> is the <em>final</em> start byte of the code unit, <code>idx_0</code> is between 0 and 5 (inclusive), and we move towards checking for overlong encoding. Just as ordinals count the number of things less than themselves, <code>idx_0</code> counts the number of content bits contained <code>start_byte</code>, occupying the least significant bits.</p>\n<p>There are six cases for anti-overlong checking, which correspond to <code>idx_0</code>&#39;s value. That may sound like a lot, but the looping gif below that I made should relax you. It demonstrates periodic behavior. Even though it shows deep</p>\n<p> code unit sections with multiple start bytes, this animation still applies for <em>all</em> code units of length <code>n &gt; 2</code>. The colored bars are based off the anatomy section image.</p>\n<pre><code>            If `idx_0 == 5` then all the mandatory content bits are contained together in the final start byte. Thus we should immediately check `start_byte` using the mask `00011111`. We then read the first non-start byte, a continuation byte which does not need overlong checking (`10xxxxxx`).\n          \n\n            Otherwise we read another continuation byte, the first non-start byte. If `idx_0 == 0` then all the mandatory content bits are contained together in this first non-start byte (`10xxxxxx`), and we use the mask `00111110` to check for overlong encoding. Else `idx_0` is between 1 and 4 (inclusive) and the mandatory content bits are straddled across the final start byte and first non-start byte. In these cases we use two masks to check for overlong encoding, which one can see in the gif above.</code></pre>\n<p>Perhaps the case of <code>idx_0 == 0</code> could be grouped in with <code>idx_0</code> being between 1 and 4, by using an empty mask to check the final start byte, in order to make the algorithm less branch-y, but this walkthrough isolates which bytes are responsible for potential overlong encoding.</p>\n<pre><code>            Given that the final start byte and first non-start byte have passed the overlong check, append them to `parsed_bytes`. Finally while the length of `parsed_bytes` is less than `n_bytes_expected`, read plain-old continuation bytes (`10xxxxxx`) and append them to `parsed_bytes`.</code></pre>\n<p>And we&#39;re done!</p>\n<h2 id=\"further-ideas\">Further Ideas</h2>\n<h2 id=\"signed-variant-zigzag-encoding\">Signed Variant: ZigZag Encoding</h2>\n<p>So far we have used UTF-8000 to encode codepoints, aka non-negative integers, aka unsigned integers. I have come up with a couple of modified interpretations of the content bits which allow us to encode the <em>entire</em> integers, aka the signed integers.</p>\n<p>We make use of a marvelous bijective mapping between the signed integers and unsigned integers called the zigzag function</p>\n<p> that remains a bijection when restricting to the respective <code>n</code>-bit ranges. We use this as a final layer at the beginning/end of the standard UTF-8000 encoding/decoding procedure.</p>\n<h3 id=\"source\">Source</h3>\n<pre><code>              ZigZag encoding from Protobuf by Google: *Protocol Buffers Documentation / Encoding*</code></pre>\n<p>Myself: This seems like the perfect extensible solution for how to encode signed integers on top of UTF-8000.</p>\n<h3 id=\"tldr-examples-2\">TLDR / Examples</h3>\n<p>The code unit structure is identical to UTF-8000. The content bits correspond to the image of the zigzag function.</p>\n<div class=\"table-wrap\"><table><thead><tr><th><code>zigzag(z)</code></th><th><code>z</code></th><th><code>UTF-8000</code></th><th></th></tr></thead><tbody><tr><td>...</td><td>...</td><td>...</td><td></td></tr><tr><td><code>124</code></td><td><code>62</code></td><td><code>01111100</code></td><td></td></tr><tr><td><code>125</code></td><td><code>-63</code></td><td><code>01111101</code></td><td></td></tr><tr><td><code>126</code></td><td><code>63</code></td><td><code>01111110</code></td><td></td></tr><tr><td><code>127</code></td><td><code>-64</code></td><td><code>01111111</code></td><td></td></tr><tr><td><code>128</code></td><td><code>64</code></td><td><code>11000010</code></td><td><code>10000000</code></td></tr><tr><td><code>129</code></td><td><code>-65</code></td><td><code>11000010</code></td><td><code>10000001</code></td></tr><tr><td><code>130</code></td><td><code>65</code></td><td><code>11000010</code></td><td><code>10000010</code></td></tr><tr><td><code>131</code></td><td><code>-66</code></td><td><code>11000010</code></td><td><code>10000011</code></td></tr><tr><td>...</td><td></td><td></td><td></td></tr></tbody></table></div>\n<h3 id=\"the-zigzag-function\">The ZigZag Function</h3>\n<p>The <code>zigzag</code> function maps from the signed integers to the unsigned integers.</p>\n<p>If <code>z ≥ 0</code> then <code>zigzag(z) = 2 * z = (z &lt;&lt; 1)</code></p>\n<p>If <code>z &lt; 0</code> then <code>zigzag(z) = -2 * z - 1 = -(z &lt;&lt; 1) - 1 = ~(z &lt;&lt; 1)</code></p>\n<div class=\"table-wrap\"><table><thead><tr><th><code>z</code></th><th><code>zigzag(z)</code></th></tr></thead><tbody><tr><td>...</td><td>...</td></tr><tr><td><code>-4</code></td><td><code>7</code></td></tr><tr><td><code>-3</code></td><td><code>5</code></td></tr><tr><td><code>-2</code></td><td><code>3</code></td></tr><tr><td><code>-1</code></td><td><code>1</code></td></tr><tr><td><code>0</code></td><td><code>0</code></td></tr><tr><td><code>1</code></td><td><code>2</code></td></tr><tr><td><code>2</code></td><td><code>4</code></td></tr><tr><td><code>3</code></td><td><code>6</code></td></tr><tr><td>...</td><td>...</td></tr></tbody></table></div>\n<div class=\"table-wrap\"><table><thead><tr><th><code>zigzag(z)</code></th><th><code>z</code></th></tr></thead><tbody><tr><td>...</td><td>...</td></tr><tr><td><code>0</code></td><td><code>0</code></td></tr><tr><td><code>1</code></td><td><code>-1</code></td></tr><tr><td><code>2</code></td><td><code>1</code></td></tr><tr><td><code>3</code></td><td><code>-2</code></td></tr><tr><td><code>4</code></td><td><code>2</code></td></tr><tr><td><code>5</code></td><td><code>-3</code></td></tr><tr><td><code>6</code></td><td><code>3</code></td></tr><tr><td><code>7</code></td><td><code>-4</code></td></tr><tr><td>...</td><td>...</td></tr></tbody></table></div>\n<pre><code>       ______________\n      /  __________  \\\n     /  /  ______  \\  \\\n    /  /  /  __  \\  \\  \\\n   /  /  /  /  \\  \\  \\  \\\n  -4 -3 -2 -1  0  1  2  3\n   \\  \\  \\  \\_____/  /  /\n    .  \\  \\_________/  /\n     .  \\_____________/\n      .\n                    </code></pre>\n<pre><code>              The ASCII art above illustrates the enumeration of the preimage of `zigzag`, showing it zigzagging between positives and negatives. This should make it clear how after `2^N` steps we have covered exactly the range `[-2^(N-1), 2^(N-1))`.</code></pre>\n<p>For example the preimage of the 7-bit unsigned range <code>[0, 128)</code> is the 7-bit signed range <code>[-64, 64)</code>.</p>\n<h4 id=\"branchless-zigzag\">Branchless ZigZag</h4>\n<p>If one is dealing with fixed-width integers, for example mapping from <code>int32_t</code> to <code>uint32_t</code>, one can create a branchless version of <code>zigzag</code>, wow! CPUs like branchless code.</p>\n<pre><code>              With this example `zigzag(z) = (z &lt;&lt; 1) ^ (z &gt;&gt; 31)`, where </code></pre>\n<p> here is the C bitwise-xor operator.\n                <code>^</code></p>\n<p>If <code>2^31 &gt; z ≥ 0</code> then <code>(z &gt;&gt; 31) = 0</code>, because we have filled the register with the highest bit of a non-negative signed number, 0. Thus <code>zigzag(z) = (z &lt;&lt; 1) ^ (z &gt;&gt; 31)</code>. Okay, nothing special?</p>\n<p>But if <code>-2^31 ≤ z &lt; 0</code> then <code>(z &gt;&gt; 31) = -1</code>, because we have filled the register with the highest bit of a negative signed number, 1. Aha, so to achieve the bitwise complement, <code>~(z &lt;&lt; 1)</code>, we can bitwise-xor with this <code>-1</code>. Thus <code>zigzag(z) = (z &lt;&lt; 1) ^ (z &gt;&gt; 31)</code>.</p>\n<h3 id=\"properties-2\">Properties</h3>\n<p>Self-synchronization, self-punctuation, and arbitrary code unit length have the same conclusion as base UTF-8000. Properties that differ are discussed below.</p>\n<h4 id=\"encoded-range\">Encoded Range</h4>\n<p>The content bit counts work the same as they do for UTF-8000. Below is a summary of the ranges of integers that the content bits encode.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>code unit length</th><th>number of content bits</th><th>minimum integer</th><th>maximum integer</th></tr></thead><tbody><tr><td><code>n = 1</code></td><td><code>7</code></td><td><code>-2 ^ (7n-1) ( = -64)</code></td><td><code>+2 ^ (7n-1) - 1 ( = +63)</code></td></tr><tr><td><code>n ≥ 2</code></td><td><code>5n+1</code></td><td><code>-2 ^ (5n)</code></td><td><code>+2 ^ (5n) - 1</code></td></tr></tbody></table></div>\n<h4 id=\"small-integers-small-code-units\">Small Integers, Small Code Units</h4>\n<p>UTF-8000 is really just a variable-width bit container with some nice properties. Provided that we obey the forbidding of overlong encoding, we can use the content bits as we please, encoding from an arbitrary alphabet to unsigned integer codewords that form the content bits.</p>\n<p>The alphabet in question for us is the signed integers, ℤ. We heuristically think of magnitude as a measure of commonness. The closer an integer is to zero, the more common it is, and thus the smaller the unsigned integer that it should be encoded as, whence the shorter the UTF-8000 code unit it occupies. This is almost common sense.</p>\n<p>We have observed that <code>zigzag</code> achieves this. Two&#39;s complements in a fixed-width register however does <em>not</em> do this, as for example in a 64-bit CPU register the number -1 is encoded as <code>111...[64]...11</code>. This is not a problem for hardware like CPUs, but we are interested in efficient encoding for transmission and storage.</p>\n<h4 id=\"modified-strcmp-3-ordering\">Modified <code>strcmp(3)</code> Ordering</h4>\n<pre><code>              Since the negative integers are interwoven (zigzagged) between the non-negative integers via `zigzag`, we lose `strcmp` ordering from UTF-8000. For example -1 &lt; 0 but `00000001` &gt; `00000000`. However, being undeterred we can supersede this fact.</code></pre>\n<p>The purpose of <code>strcmp(s1, s2)</code> with respect to UTF-8000 is to quickly compare code units <code>s1</code> and <code>s2</code> as a proxy for comparing their contained codepoints, without having to actually decode the code units. For this modified version of UTF-8000 we wish to create a quick proxy for comparing the contained signed integers.</p>\n<p>The only variants of UTF-8000 that can make exact use of <code>strcmp</code> are those whose content bits encode letters from a totally-ordered alphabet, for which there exists an order-preserving bijection between that alphabet and the unsigned integers. Since the unsigned integers has a minimum element, 0, and the signed integers (our alphabet) does not have a minimum element, no such bijection exists.</p>\n<p>We know that <code>zigzag(z)</code> breaks nicely into two cases, non-negative signed integers and negative signed integers. We also know that order-preserving bijections <em>do</em> exist between 0) non-negative signed integers and the even unsigned integers, and 1) negative signed integers and the odd unsigned integers. We initially break our new <code>strcmpsigned(s1, s2)</code> function into four cases depending on the final bit of each code unit, which we know is a content bit and indicates whether the stored unsigned integer is even or odd. We can obtain these bits via <code>b1 = c1 &amp; 1</code> and <code>b2 = c2 &amp; 1</code> where <code>c1</code> and <code>c2</code> are the final bytes of the respective code units:</p>\n<div class=\"table-wrap\"><table><thead><tr><th><code>b1</code></th><th><code>b2</code></th><th><code>comment</code></th><th><code>b2 - b1</code></th><th><code>1 - b1 - b2</code></th></tr></thead><tbody><tr><td><code>0</code></td><td><code>0</code></td><td><code>s1 ? s2</code></td><td><code>0</code></td><td><code>1</code></td></tr><tr><td><code>0</code></td><td><code>1</code></td><td><code>s1 &gt; s2</code></td><td><code>1</code></td><td><code>0</code></td></tr><tr><td><code>1</code></td><td><code>0</code></td><td><code>s1 &lt; s2</code></td><td><code>-1</code></td><td><code>0</code></td></tr><tr><td><code>1</code></td><td><code>1</code></td><td><code>s1 ? s2</code></td><td><code>0</code></td><td><code>-1</code></td></tr></tbody></table></div>\n<p>If <code>b2 - b1</code> is non-zero then <code>strcmpsigned</code> can return that, as we are effectively comparing two signed integers of a different sign. Otherwise <code>(1 - b1 - b2) * strcmp(s1, s2)</code> effectively compares two integers of the same sign. We could therefore write this as:</p>\n<pre><code>              `strcmpsigned(s1, s2) = (b2 - b1) ? (b2 - b1) : (1 - b1 - b2) * strcmp(s1, s2)`</code></pre>\n<p>That&#39;s pretty succinct! I have also assumed that <code>strcmp</code> is just the simple <code>{-1, 0, +1}</code> version.</p>\n<h3 id=\"verdict\">Verdict</h3>\n<pre><code>              I like it! I&#39;ll add it to the reference implementation when I get chance. It will most likely be a flag </code></pre>\n<p> for <code>-z</code>\nzigzag</p>\n<p> used like <code>$ utf-8000 info -z -- -67</code> showing <code>11000010</code> <code>10000101</code>.</p>\n<p>This zigzag variant has a big advantage over the rejected two&#39;s complement signed variant in that we don&#39;t need to change how we do overlong checking from the standard UTF-8000 method. This means that we can effectively separate out into layers</p>\n<p>: 1) the overlong checking of code units and the extracting of their content bits, from 2) the further decoding of the content bits eg to a signed integer.</p>\n<p>The only external metadata needed when decoding a stream of UTF-8000 bytes is is this unsigned or signed?</p>\n<p>. This is no different to decoding a stream of raw bytes, or inspecting fixed-width integers stored in two&#39;s complement form in a CPU register: it&#39;s up to the programmer&#39;s use-case to know whether signed or unsigned is expected.</p>\n<h2 id=\"utf-16k\">UTF-16K</h2>\n<p>Could we apply some techniques from this document to also extend UTF-16? Yes, but at the cost of forbidding more codepoints from being encoded similar to the forbidden surrogate range <code>U+D800</code> to <code>U+DFFF</code>.</p>\n<p>UTF-16 is a bit messy in its existing two-byte and four-byte form, but we can clean this up in a mostly forwards-compatible manner by using ASCVI-on-UTF-16. We assume the use of big-endian UTF-16 in this section.</p>\n<pre><code>              The high (`110110`) and low (`110111`) surrogate prefixes are highlighted in bright pink.</code></pre>\n<p>For normal 20-content-bit surrogate pair UTF-16, the upper four content bits of high surrogates encode which Unicode Plane (collection of <code>2^16 = 64k</code> codepoints) that the code unit&#39;s content bits belong to. These bits are highlighted in bright crimson.</p>\n<p>For multi-surrogate-pair UTF-16K we highlight only the upper three of these plane bits, the erstwhile fourth being a self-synchronization bit.</p>\n<h3 id=\"source-2\">Source</h3>\n<p>Myself: Realizing that I can challenge the UTF-16 extensions proposed by UCS-X.</p>\n<h3 id=\"tldr-examples-3\">TLDR / Examples</h3>\n<p>| 2 | <code>xxxxxxxx xxxxxxxx</code> |  |  |  |  |  |  | \n| 4 | <code>110110xx xxxxxxxx</code> | <code>110111xx xxxxxxxx</code> |  |  |  |  |  | \n| 8 | <code>11011010 000xxxxx</code> | <code>110111xx xxxxxxxx</code> | <code>11011010 01xxxxxx</code> | <code>110111xx xxxxxxxx</code> |  |  |  | \n| 12 | <code>11011010 0010xxxx</code> | <code>110111xx xxxxxxxx</code> | <code>11011010 01xxxxxx</code> | <code>110111xx xxxxxxxx</code> | <code>11011010 01xxxxxx</code> | <code>110111xx xxxxxxxx</code> |  | \n| 16 | <code>11011010 00110xxx</code> | <code>110111xx xxxxxxxx</code> | <code>11011010 01xxxxxx</code> | <code>110111xx xxxxxxxx</code> | <code>11011010 01xxxxxx</code> | <code>110111xx xxxxxxxx</code> | ... | \n| 20 | <code>11011010 001110xx</code> | <code>110111xx xxxxxxxx</code> | <code>11011010 01xxxxxx</code> | <code>110111xx xxxxxxxx</code> | <code>11011010 01xxxxxx</code> | <code>110111xx xxxxxxxx</code> | ... | \n| ... |  |  |  |  |  |  |  | \n| 76 | <code>11011010 00111111</code> | <code>11011111 11111111</code> | <code>11011010 0110xxxx</code> | <code>110111xx xxxxxxxx</code> | <code>11011010 01xxxxxx</code> | <code>110111xx xxxxxxxx</code> | ... | \n| ... |  |  |  |  |  |  |  | \n| 144 | <code>11011010 00111111</code> | <code>11011111 11111111</code> | <code>11011010 01111111</code> | <code>11011111 11111111</code> | <code>11011010 01110xxx</code> | <code>110111xx xxxxxxxx</code> | ... | \n| ... |  |  |  |  |  |  |  | </p>\n<h3 id=\"properties-3\">Properties</h3>\n<pre><code>              Two-byte UTF-16 is just the raw binary form of any 16-bit codepoint, except for the surrogate</code></pre>\n<p> range <code>U+D800</code> to <code>U+DFFF</code> of size 2048 which any Unicode encoding (UTF-8, UTF-16, UTF-32) is forbidden to encode. The reason for this exclusion is because otherwise we could not distinguish <code>110110xx xxxxxxxx</code> from <code>110110xx xxxxxxxx</code> and <code>110111xx xxxxxxxx</code> from <code>110111xx xxxxxxxx</code> which you&#39;ll read about below.</p>\n<pre><code>              Four-byte UTF-16 is two surrogate codepoints stuck next to each other, one high</code></pre>\n<p> in the range <code>U+D800</code> to <code>U+DBFF</code> and then one low</p>\n<p> in the range <code>U+DC00</code> to <code>U+DFFF</code>. The real codepoint that they encode is 0x10000 added to the 20 binary digit number contained in the content bits within. For example <code>11011000 00001000</code> <code>11011100 00101101</code> contains 0x0202D, which then has 0x10000 added to it, to give 0x1202D. <code>U+1202D</code> is 𒀭</p>\n<p>, a Mesopotamian Dingir.</p>\n<pre><code>              To expand UTF-16 indefinitely instead of being stuck with 0x110000 (1,114,112) codepoints, we employ ASCVI inside of UTF-16 surrogate pairs. UTF-8 was able to expand from 7-bit ASCII without any trouble because bytes with the highest bit set were undefined. In UTF-16 however every possible value that the content bits can take defines a codepoint. A smart choice, so that we do not interfere with already assigned codepoints, and so that we do not malapportion too many pre-existing unassigned codepoints for UTF-16K, and for there to be content bit count parity with UTF-8000, is to constrain ourselves to two unassigned planes, e.g. Planes 9 and 10. These planes are nice as the surrogate pairs take the form `11011010 0xxxxxx` `110111xx xxxxxxxx`. The codepoints `U+90000` to `U+AFFFF` are to be forbidden from being encoded, just as the 2048 surrogates of the Basic Multilingual Plane are.</code></pre>\n<p>We use these four-byte surrogate pair containers as the quantum for UTF-16K, which uses two or more of these quanta to extend from UTF-16. The content bits encode the codepoint&#39;s binary representation, <em>without</em> adding on 0x10000 for the sake of simplicity, similar to UTF-8000.</p>\n<h4 id=\"bit-counts-2\">Bit Counts</h4>\n<div class=\"table-wrap\"><table><thead><tr><th>number of bytes</th><th>number of content bits</th><th>number of mandatory content bits</th></tr></thead><tbody><tr><td><code>n = 2</code></td><td><code>16</code></td><td><code>0</code></td></tr><tr><td><code>n = 4</code></td><td><code>20</code></td><td><code>0</code></td></tr><tr><td><code>n = 8  ; k = 2</code></td><td><code>15k + 1 ( = 31)</code></td><td><code>10, or codepoint ≥ 0x110000</code></td></tr><tr><td><code>n = 4k ; k ≥ 3</code></td><td><code>15k + 1</code></td><td><code>15</code></td></tr></tbody></table></div>\n<p>Overlong checking is complicated for the jump from 4 bytes to 8 bytes, because of the 0x10000 that is added to the content bits. Either one of the highest 10 bits has a non-zero bit, or the <code>2^20</code> bit is active and one of the <code>2^m</code> bits is active with <code>16 ≤ m ≤ 19</code>.</p>\n<p>One may notice when enumerating <code>15k + 1</code>, the number of content bits for <code>k</code>-surrogate-pair UTF-16K, that these values overlap predictably with UTF-8000&#39;s number of content bits given by <code>5n + 1</code>. This is the result of a deliberate choice to use two planes for UTF-16K instead of e.g. one plane, or half a plane etc.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>number of UTF-16K surrogate pairs</th><th>number of UTF-8K bytes</th><th>number of content bits</th></tr></thead><tbody><tr><td><code>2</code></td><td><code>6</code></td><td><code>31</code></td></tr><tr><td><code>3</code></td><td><code>9</code></td><td><code>46</code></td></tr><tr><td><code>4</code></td><td><code>12</code></td><td><code>61</code></td></tr><tr><td><code>5</code></td><td><code>15</code></td><td><code>76</code></td></tr><tr><td>...</td><td>...</td><td>...</td></tr><tr><td><code>17</code></td><td><code>51</code></td><td><code>256</code></td></tr><tr><td>...</td><td>...</td><td>...</td></tr></tbody></table></div>\n<p>This means that UTF-8K and UTF-16K can be expanded in parallel in a predictable way, such that all possible codepoints from an expansion are permitted. This is in contrast to how 4-byte UTF-8 does not allow use of all <code>2^21</code> codepoints, but rather artificially restricts to <code>2^16 + 2^20</code> for parity with UTF-16K.</p>\n<p>The number of bytes in such UTF-16K code units is always <code>4/3</code> that of an equivalent UTF-8K code unit. The reciprocal of this, <code>3/4</code>, ends up as the scaling factor of the information rate limit from UTF-8K to UTF-16K. It is nice that this ratio is independent of code unit length.</p>\n<h4 id=\"information-rate-2\">Information Rate</h4>\n<p>For 2-byte UTF-16 this is technically <code>16 / 16 = 100%</code>, ignoring the forbidden surrogate range.</p>\n<p>For 4-byte UTF-16 this is <code>20 / 32 = 62.5%</code>, again ignoring the forbidden UTF-16K planes 9 and 10. This looks slightly worse than UTF-8&#39;s <code>21 / 32</code>, but do bear in mind that UTF-8 is also restrained to UTF-16&#39;s upper limit.</p>\n<p>For beyond four bytes this is <code>(15k+1) / (4k*8) = 15/32 + 1/(32k)</code> which approaches <code>15/32 = 46.875%</code>, which is okay. As predicted above, this is <code>3/4</code> times the information rate limit of UTF-8000. <code>3/4 * 5/8 = 15/32</code>.</p>\n<p>Below is a comparison of the efficiencies of UTF-8 and UTF-16.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>start</th><th>end</th><th>range size</th><th>number of UTF-8 bytes</th><th>number of UTF-16 bytes</th></tr></thead><tbody><tr><td><code>U+0000</code></td><td><code>U+007F</code></td><td><code>0x80 = 128</code></td><td><code>1 (ASCII)</code></td><td><code>2</code></td></tr><tr><td><code>U+0080</code></td><td><code>U+07FF</code></td><td><code>0x780 = 1,920</code></td><td><code>2</code></td><td><code>2</code></td></tr><tr><td><code>U+0800</code></td><td><code>U+FFFF</code></td><td><code>0xF800 = 63,488</code></td><td><code>3</code></td><td><code>2</code></td></tr><tr><td><code>U+10000</code></td><td><code>U+10FFFF</code></td><td><code>0x100000 = 1,048,576</code></td><td><code>4</code></td><td><code>4</code></td></tr></tbody></table></div>\n<p>UTF-16 is more efficient than UTF-8 only at encoding <code>U+0800</code> to <code>U+FFFF</code>, aka the three-byte UTF-8 range that UTF-16 encodes using two bytes.</p>\n<p>For ASCII, and beyond <code>U+10FFFF</code>, UTF-8000 is far more efficient (and less ugly) than UTF-16K.</p>\n<h4 id=\"self-synchronization-2\">Self-Synchronization</h4>\n<p>UTF-16K does not have self-synchronization at the byte level because UTF-16 does not. If one experiences a single missing byte then potentially the whole stream becomes corrupted.</p>\n<pre><code>              At the two-byte level UTF-16 has self-synchronization which UTF-16K inherits. Non-surrogate codepoints are quantum; they are to UTF-16 as ASCII is to UTF-8. Surrogate pairs provide self-synchronization with their `110110` and `110111` high and low surrogate prefixes.\n            \n\n              UTF-16K goes even deeper, using multiple surrogate pairs that shadow Planes 9 and 10. Within the surrogate pair self-synchronization level, within the high surrogates used to encode UTF-16K, `11011010 0xxxxxx`, ASCVI is employed, whose leading bit provides self-synchronization, with `11011010 00` and `11011010 01`.</code></pre>\n<p>Therefore overall UTF-16K exhibits self-synchronization at the two-byte level, like UTF-16.</p>\n<h4 id=\"self-punctuation-2\">Self-Punctuation</h4>\n<p>Inherited from ASCVI.</p>\n<h4 id=\"compatibility\">Compatibility</h4>\n<p>UTF-16K forbids Planes 9 and 10 of Unicode, because it has no way to encode those codepoints, instead repurposing the surrogate pairs erstwhile required to encode Planes 9 and 10 for the purpose of encoding codepoints beyond 0x110000. An important question to ask regarding forwards compatibility is what does an existing UTF-16 parser do if it meets a UTF-16K code unit?</p>\n<p>.</p>\n<p>In short it&#39;s Plane-9-or-10-garbage-in Plane-9-or-10-garbage-out. Each surrogate pair used in encoding a UTF-16K codepoint beyond 0x110000 would be parsed separately as though it belongs to Plane 9 or 10, but with no <em>syntactic</em> issues. <em>Semantically</em> however this would cause issues with logical character (codepoint) counts that would count each surrogate pair as a separate character, rather than contributing towards a single character.</p>\n<p>I think that this is a better solution than UCS-X&#39;s UTF-G-16 which breaks syntactic compatibility with UTF-16 by repurposing low surrogates as leading units</p>\n<p> for UTF-G-16 6-byte code units. One could argue that UTF-G-16 is better because those bytes could be replaced with a single replacement character �</p>\n<p> though I&#39;m not convinced, as for example the default behavior of Python&#39;s <code>bytes.decode</code> function is <code>&#39;strict&#39;</code>, which raises an exception, not <code>&#39;replace&#39;</code> which produces replacement characters. UTF-G-16 also has flawed error handling behavior as discussed in the feedback emails, arising from UTF-G-16&#39;s self-synchronization requiring a context-dependent interpretation of low surrogates to determine if they are leading</p>\n<p> or trailing</p>\n<p>, whereas UTF-16K&#39;s self-synchronization is context-independent by using a disjoint union of planes 9 and 10.</p>\n<p>The requirement to extend the list of codepoints that all of UTF-8, UTF-16, and UTF-32 are forbidden from encoding, to include Planes 9 and 10 or elsewhere, would not be an easily negotiated feat. We would be banning an extra <code>2/17 = 11.8%</code> of pre-existing codepoints. One may notice that this situation of forbidding pre-existing codepoints is a similar situation to back when 2-byte UTF-16 extended to 4-byte UTF-16. Would we ever have to ban codepoints in pre-existing ranges again after this UTF-16K extension? No, as UTF-8K and UTF-16K are infinitely extensible.</p>\n<h3 id=\"verdict-2\">Verdict</h3>\n<pre><code>              The immature part of me says let UTF-16 decay and die as the short-sighted, legacy, Windows, </code></pre>\n<p>. But it will be around for a while, with several uses, like the Joliet Filesystem for my beloved Arch Linux ISOs grrr.\n                <code>wchar_t</code>, non-self-synchronizing-at-the-byte-level, +0x10000, garbage that it is</p>\n<p>The main reason I wrote this section was to provide an alternative to UCS-X, so that we can use the same(ish) style as UTF-8000, and lest UCS-X or an even uglier idea come along.</p>\n<p>I predict that the Unicode Consortium would heavily push back on the idea of having to ban more codepoints, <code>U+90000</code> to <code>U+AFFFF</code>. No matter how one plans to extend UTF-16, it requires either forbidding some codepoints, or changing the syntax, either of which is a breaking change.</p>\n<p>For our modern times UTF-8 is undoubtedly the way to go, and by the time that we need to extend to UTF-8000, I would hope that UTF-16 and all other encodings belong in a museum, and we can therefore extend UTF-8 without worrying about compatibility with the others.</p>\n<h3 id=\"reference-implementation\">Reference Implementation</h3>\n<p>Available! See below.</p>\n<h2 id=\"utf-32k\">UTF-32K</h2>\n<p>In the same manner that ASCII is extended by UTF-8 and UTF-8000, UTF-32 could also be extended to be a multi-\nbyte</p>\n<p> (32-bit chunk) encoding scheme.</p>\n<p>Do we really want this though? Is UTF-32 meant to be variable-width, or is it meant to represent the raw codepoint, decoded and stored in memory as a fixed-width integer? I&#39;ve written this section to demonstrate that ASCVI can be applied to a quantum as small as 3 bits (seriously lol), or large like 32 bits.</p>\n<h3 id=\"source-3\">Source</h3>\n<p>Myself: It seemed obvious how this follows from UTF-8000.</p>\n<h3 id=\"tldr-examples-4\">TLDR / Examples</h3>\n<p>We could extend UTF-32 either in the style of UTF-8000, treating the one-\nbyte</p>\n<p> code units as a special case occupying the lower 31 bits...</p>\n<p>| 1 | <code>0xxxxxxx_xxxxxxxx_xxxxxxxx_xxxxxxxx</code> |  | \n| 2 | <code>110xxxxx_xxxxxxxx_xxxxxxxx_xxxxxxxx</code> | <code>10xxxxxx_xxxxxxxx_xxxxxxxx_xxxxxxxx</code> | \n| ... |  |  | </p>\n<p>...or in the style of ASCVI, with no special cases and the one-\nbyte</p>\n<p> code units occupying the lower 30 bits.</p>\n<p>| 1 | <code>00xxxxxx_xxxxxxxx_xxxxxxxx_xxxxxxxx</code> |  | \n| 2 | <code>010xxxxx_xxxxxxxx_xxxxxxxx_xxxxxxxx</code> | <code>1xxxxxxx_xxxxxxxx_xxxxxxxx_xxxxxxxx</code> | \n| ... |  |  | </p>\n<h3 id=\"properties-4\">Properties</h3>\n<p>Mutatis mutandis, the properties of UTF-8000 and ASCVI apply. We only make further remarks on a couple of properties.</p>\n<h4 id=\"bit-counts-3\">Bit Counts</h4>\n<p>Predictable like ASCVI.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>number of bytes</th><th>number of content bits</th><th>number of mandatory content bits</th></tr></thead><tbody><tr><td><code>n = 4</code></td><td><code>30</code></td><td><code>0</code></td></tr><tr><td><code>n = 4k ; k ≥ 2</code></td><td><code>30k</code></td><td><code>30</code></td></tr></tbody></table></div>\n<h4 id=\"information-rate-3\">Information Rate</h4>\n<p>Whilst the ASCVI-style UTF-32K has an information rate of <code>30 / 32 = 93.75%</code>, it is very inefficient for low-value codepoints, with the highest bytes most often being zeros. UTF-8000 has a much finer telescopic expansion mechanism at the byte level, compared to UTF-32 at a four-byte level.</p>\n<h4 id=\"self-synchronization-3\">Self-Synchronization</h4>\n<p>UTF-32 is not self-synchronizing at the byte level, and UTF-32K inherits this weakness. UTF-32 is only self-synchronizing at the four-byte level, similar to how UTF-16 is only self-synchronizing at the two-byte level. UTF-32K maintains self-synchronization at the four-byte level.</p>\n<h4 id=\"endianness\">Endianness</h4>\n<p>Like UTF-16, and unlike UTF-8 and UTF-8000, UTF-32 has endianness, its quantum being a whopping four bytes.</p>\n<h3 id=\"verdict-3\">Verdict</h3>\n<p>Not our greatest priority.</p>\n<p>UTF-32 is hardly ever used for transmission or storage due to its inefficiency and endianness.</p>\n<p>As I wrote in the intro, UTF-32&#39;s main use is as a <strong>non</strong>-variable-width container, for when one decodes UTF-8 or UTF-16 to <code>int32_t</code> integers (UTF-32) for use inside a program. UTF-32K would be an anti-pattern / counterproductive.</p>\n<h2 id=\"rejected-alternatives\">Rejected Alternatives</h2>\n<p>Although UTF-8000 extends naturally</p>\n<p> from UTF-8, is it still the best approach? Are there any better alternatives that engineer</p>\n<p> an extension from UTF-8, just as UTF-8 engineers an extension from ASCII?</p>\n<p>We rule out a few alternatives in this section. It&#39;s good to document the suboptimal solutions (and outright failures) so that we can work towards success. I&#39;ve done that plenty of times with my own ideas don&#39;t worry! Feel satisfied in having at least made an attempt.</p>\n<h2 id=\"ascvi\">ASCVI</h2>\n<p>What if ASCII were only six-bit instead of seven-bit? Would this make extending to multi-byte code units more pleasant?</p>\n<h3 id=\"source-4\">Source</h3>\n<p>Myself: The intuitive derivation section of UTF-8000.</p>\n<h3 id=\"tldr-examples-5\">TLDR / Examples</h3>\n<p>| 1 | <code>00xxxxxx</code> |  |  |  |  |  | \n| 2 | <code>010xxxxx</code> | <code>1xxxxxxx</code> |  |  |  |  | \n| 3 | <code>0110xxxx</code> | <code>1xxxxxxx</code> | <code>1xxxxxxx</code> |  |  |  | \n| ... |  |  |  |  |  |  | \n| 17 | <code>01111111</code> | <code>11111111</code> | <code>1110xxxx</code> | <code>1xxxxxxx</code> | ... | <code>1xxxxxxx</code> | \n| ... |  |  |  |  |  |  | </p>\n<h3 id=\"properties-5\">Properties</h3>\n<p>Self-synchronization, self-punctuation, <code>strcmp</code> ordering, BOM support, and arbitrary code unit length have the same conclusion as UTF-8000. Properties that differ are discussed below.</p>\n<h4 id=\"bit-counts-4\">Bit Counts</h4>\n<p>The number of content bits and mandatory content bits are even more predictable than those of UTF-8.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>code unit length</th><th>number of content bits</th><th>number of mandatory content bits</th></tr></thead><tbody><tr><td><code>n = 1</code></td><td><code>6n ( = 6)</code></td><td><code>0</code></td></tr><tr><td><code>n &gt; 1</code></td><td><code>6n</code></td><td><code>6</code></td></tr></tbody></table></div>\n<pre><code>              This is because one-byte code units are not special. They use the same `0` self-synchronization prefix as any first byte. The number 6 arises from every subsequent continuation byte adding on 7 more content bits, minus 1 for the longer start bits sequence.</code></pre>\n<p>Consequently the number of content bits stored in an <code>n</code> byte code unit is never a power of two, unlike with UTF-8000. This is because 6, containing 3 in its prime factorization, cannot divide into a power of two. This is not a terrible defect, but we do like powers of two.</p>\n<h4 id=\"information-rate-4\">Information Rate</h4>\n<p>ASCVI&#39;s information rate is <code>6n / 8n = 6 / 8 = 75%</code>. This a constant independent of the length of the code unit.</p>\n<p>For one-byte code units UTF-8000 (ASCII) is more efficient and versatile, storing double the number of codepoints and having an information rate of <code>87.5%</code>.</p>\n<p>For multi-byte code units ASCVI is more efficient, with UTF-8000&#39;s information rate tending downwards towards <code>62.5%</code>.</p>\n<p>Even if the US English alphabet had its 52 letters cut down to eg 27 Hebrew glyphs, or no letters at all, one would struggle to create a practical set of 64 glyphs for single-byte ASCVI. The tradeoff of ASCII being seven-bit, at the slight detriment of the information rate of multi-byte code units, seems worth it.</p>\n<h4 id=\"byte-map-2\">Byte Map</h4>\n<div class=\"table-wrap\"><table><thead><tr><th></th><th>0</th><th>1</th><th>2</th><th>3</th><th>4</th><th>5</th><th>6</th><th>7</th><th>8</th><th>9</th><th>A</th><th>B</th><th>C</th><th>D</th><th>E</th><th>F</th></tr></thead><tbody><tr><td>0</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td></tr><tr><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td></tr><tr><td>2</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td></tr><tr><td>3</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td></tr><tr><td>4</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td></tr><tr><td>5</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td><td>2</td></tr><tr><td>6</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td><td>3</td></tr><tr><td>7</td><td>4</td><td>4</td><td>4</td><td>4</td><td>4</td><td>4</td><td>4</td><td>4</td><td>5</td><td>5</td><td>5</td><td>5</td><td>6</td><td>6</td><td>7</td><td>8+</td></tr><tr><td>8</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>9</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>A</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>B</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>C</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>D</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>E</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>F</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr></tbody></table></div>\n<p>Look at that beautiful geometric series layout. Four rows for <code>1</code>, two rows for <code>2</code>, one row for <code>3</code>, half a row for <code>4</code> etc...</p>\n<p>Unlike UTF-8000 which can never use the bytes 0xC0 and 0xC1, ASCVI uses all 256 possible bytes. Which of these is the advantageous behavior depends on whether one wants to do something extraordinary with those two bytes, or one wants to dissuade their use in chicanery.</p>\n<h3 id=\"intuitive-derivation-2\">Intuitive Derivation</h3>\n<p>See the intuitive derivation of UTF-8000 for how one might come up with this.</p>\n<h3 id=\"naming\">Naming</h3>\n<p>The II</p>\n<p> in ASCII reminds me of the Roman Numerals VII</p>\n<p> for seven, and ASCII is a seven-bit code. Therefore for this six-bit code we choose to use the Roman Numerals for six, VI</p>\n<p>, and name it ASCVI</p>\n<p>.</p>\n<h3 id=\"verdict-4\">Verdict</h3>\n<p>7 bit ASCII, and UTF-8 that extends it, are very well established. One-byte ASCVI is inferior to the flexibility of ASCII, albeit this contributes to UTF-8 having a slightly lower information rate for multi-byte code units. I do not yearn for an alternate universe, or a fresh start of text encoding standards, where ASCII is six-bit instead of seven.</p>\n<p>That being said, ASCVI is by no means <em>inherently</em> flawed, and we can make use of it in UTF-16K and UTF-32K. In a sense ASCVI is the Platonic Form of UTF-8.</p>\n<h2 id=\"signed-variant-two-s-complement\">Signed Variant: Two&#39;s Complement</h2>\n<p>One might be shocked to find our beloved two&#39;s complement representation of the signed integers present in the rejected ideas section. This is not about one&#39;s personal taste in representing signed integers, but technological extensibility, with which the zigzag signed variant far outshines this two&#39;s complement variant.</p>\n<h3 id=\"source-5\">Source</h3>\n<p>Myself: It seemed like a good(ish) idea until I realized that the zigzag variant is better.</p>\n<h3 id=\"tldr-examples-6\">TLDR / Examples</h3>\n<p>We treat the <code>n</code> content bits of a code unit as a two&#39;s complement signed form, where the highest bit no longer has value <code>2 ^ (n-1)</code> but rather <code>-2 ^ (n-1)</code>.</p>\n<p>| ASCII |  |  |  |  |  | \n| 1 | <code>0xxxxxxx</code> |  |  |  |  | \n| UTF-8 |  |  |  |  |  | \n| 2 | <code>110xxxxx</code> | <code>10xxxxxx</code> |  |  |  | \n| 3 | <code>1110xxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  | \n| 4 | <code>11110xxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  | \n| UTF-8000 |  |  |  |  |  | \n| 5 | <code>111110xx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | \n| ... |  |  |  |  |  | </p>\n<p>The difference in code unit layout from UTF-8000 is that the mandatory content bits are downshifted by one bit.</p>\n<h3 id=\"properties-6\">Properties</h3>\n<h4 id=\"overlong-encoding-checking\">Overlong Encoding Checking</h4>\n<p>Preventing overlong encoding requires checking that the content bits of an <code>n</code>-byte code unit encode an integer <code>z</code> from the range <code>[ -2^(5n), +2^(5n) ) \\ [ -2^(5(n-1)), +2^(5(n-1)) )</code>. This may look complicated, but we can break it down into two cases:</p>\n<pre><code>              For `z ≥ 0`, the highest content bit of both `n`-byte and `n-1`-byte code units is `0`. There must be at least one `1` bit in the following bits, which are the mandatory content bits.\n            \n\n              For `z &lt; 0`, the highest content bit of both `n`-byte and `n-1`-byte code units is `1`. This is to make `z` negative, using the `-2^(5n)` bit. There must be at least one `0` bit in the following bits, which are the mandatory content bits. Otherwise if all of these bits were ones, then `z` would be at least `-2^(5(n-1))`. For example the overlong encoding\n              `11011111`\n              `10000000` encodes `z = -64` which fits into a 1-byte code unit. Encoding integers below `-64` requires subtracting from these content bits, which sets at least one of the mandatory content bits to zero.\n            \n\n              UTF-8000 never uses the bytes 0xC0 and 0xC1, which is explained in the glossary section for overlong encoding. Slightly differently, this two&#39;s complement signed variant never uses the bytes 0xC0 (`11000000`) or 0xDF (`11011111`).</code></pre>\n<h4 id=\"no-strcmp-3-ordering\">No <code>strcmp(3)</code> Ordering</h4>\n<pre><code>              Since negative integers set the highest content bit to `1`, we lose `strcmp` ordering from UTF-8000. For example -1 &lt; 0 but `01111111` &gt; `00000000`.</code></pre>\n<h3 id=\"verdict-5\">Verdict</h3>\n<pre><code>              A big downside of this two&#39;s complement signed variant is that its anti-overlong checking mechanism differs from that of UTF-8000 because of the 1-bit-downshifted position of the mandatory content bits. Consequently it is not possible to agnostically decode a stream of these bytes as though they were UTF-8000 bytes. For example, a 0xC1 byte is valid in this variant as `11000001`, but is invalid in UTF-8000 as `11000001`.</code></pre>\n<p>What if we tried to redeem this variant by proposing to move the negative</p>\n<p> bit, the <code>-2^(5n)</code> bit, to the stable tail end of the code unit, rather than it being at the ever-expanding head of the code unit? This could hopefully mean that we would not have to change the anti-overlong checking mechanism from UTF-8000. The long and short is that we would perchance intuitively reinvent the zigzag signed variant from first principles, which indeed leftwards bitshifts by one the signed integer that it encodes, using the lowest bit as the negative</p>\n<p> bit. Our redemption is found there.</p>\n<h2 id=\"utf-infinity\">UTF-Infinity</h2>\n<p>What if we naively roll the start bits</p>\n<p> over into further bytes?</p>\n<h3 id=\"source-6\">Source</h3>\n<pre><code>              Mashpoe on YouTube: *Expanding the UTF-8 Character Set to Infinity*</code></pre>\n<h3 id=\"tldr-examples-7\">TLDR / Examples</h3>\n<p>| ... |  |  |  |  |  |  | \n| 7 | <code>11111110</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | ... | <code>10xxxxxx</code> |  | \n| 8 | <code>11111111</code> | <code>0xxxxxxx</code> | <code>10xxxxxx</code> | ... | <code>10xxxxxx</code> |  | \n| 9 | <code>11111111</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | ... | <code>10xxxxxx</code> |  | \n| 10 | <code>11111111</code> | <code>110xxxxx</code> | <code>10xxxxxx</code> | ... | <code>10xxxxxx</code> |  | \n| ... |  |  |  |  |  |  | \n| 15 | <code>11111111</code> | <code>11111110</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | ... | <code>10xxxxxx</code> | \n| 16 | <code>11111111</code> | <code>11111111</code> | <code>0xxxxxxx</code> | <code>10xxxxxx</code> | ... | <code>10xxxxxx</code> | \n| 17 | <code>11111111</code> | <code>11111111</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | ... | <code>10xxxxxx</code> | \n| 18 | <code>11111111</code> | <code>11111111</code> | <code>110xxxxx</code> | <code>10xxxxxx</code> | ... | <code>10xxxxxx</code> | \n| ... |  |  |  |  |  |  | </p>\n<h3 id=\"properties-7\">Properties</h3>\n<h4 id=\"bit-counts-5\">Bit Counts</h4>\n<p>Effectively, every jump from <code>8k-1</code>-byte code units to <code>8k</code>-byte code units the encoding inserts another blank byte after the start bytes, by which 8 minus 1 equals 7 bits of content are gained in an ASCII-looking byte, instead of appending a continuation byte by which 6 minus 1 equals 5 bits of content are gained.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>code unit length</th><th>number of content bits</th><th>number of mandatory content bits</th></tr></thead><tbody><tr><td><code>n = 1</code></td><td><code>7</code></td><td><code>0</code></td></tr><tr><td><code>n ≠ 8k</code></td><td><code>5n+1 + 2⌊n/8⌋</code></td><td><code>5</code></td></tr><tr><td><code>n = 8k</code></td><td><code>5n+1 + 2⌊n/8⌋</code></td><td><code>7</code></td></tr></tbody></table></div>\n<h4 id=\"information-rate-5\">Information Rate</h4>\n<p>For an <code>n</code>-byte code unit the information rate is UTF-8000&#39;s information rate plus <code>2⌊n/8⌋ / (8n)</code>.</p>\n<p>I&#39;m not working it out fully, but I can tell that this leads to a sawtooth-y profile as a graph of information rate against <code>n</code>. Therefore, counterintuitively, longer code units can have better efficiency than shorter ones.</p>\n<h4 id=\"no-self-synchronization\">No Self-Synchronization</h4>\n<pre><code>              In the 8-byte code unit example, there is no way to distinguish the second byte `0xxxxxxx` from an ASCII byte `0xxxxxxx`. This generalizes to `8n`-byte code units.\n            \n\n              In the 15-byte code unit example, there is no way to distinguish the second byte, `11111110` from the first byte of a 7-byte code unit. This generalizes to `8n-1`-byte code units.\n            \n\n              In the 16-byte code unit example, there is no way to distinguish the second byte, `11111111` from the first byte of an 8-byte code unit. This generalizes such that if one seeks to any `11111111` byte, one has no idea if this is the first byte of a code unit or not.</code></pre>\n<p>This list is non-exhaustive.</p>\n<h4 id=\"self-punctuation-3\">Self-Punctuation</h4>\n<p>This is the property that Mashpoe clearly prioritized preserving, however the approach was too myopic and did not lead to preserving other properties of interest.</p>\n<h4 id=\"patented\">Patented</h4>\n<pre><code>            Mashpoe jokes (?) in the video that he owns the patent to this encoding scheme.</code></pre>\n<p>Regardless of whether he is joking or not, I nonetheless find it reprehensible that someone <em>could</em> (at least try to) copyright / patent the correct way to extend UTF-8. It would be like trying to copyright the right solution to a mathematics equation, or a prime number! Therefore I am being quite loud in the copylefting of UTF-8000 in the licensing section. Everyone benefits from shared, free-as-in-freedom, open ideas.</p>\n<h3 id=\"verdict-6\">Verdict</h3>\n<p>The loss of self-synchronization is a fatal detriment.</p>\n<p>The formula for the number of content bits has predictable but irritable jumps, which lead to counterintuitive information rates.</p>\n<h2 id=\"perl-utf8\">Perl utf8</h2>\n<pre><code>              Use *up to* 7 bytes to encode up to 36 bits of information in the sane way, in order to encode 32-bit integers (and a little beyond). To encode 64-bit integers, use a special-case *fixed* 13-byte code unit starting with `11111111`.</code></pre>\n<p>Since the <code>n</code>-byte code units with <code>n &lt; 8</code> are the same as UTF-8000 we shall mostly only discuss the 13-byte code units.</p>\n<h3 id=\"source-7\">Source</h3>\n<pre><code>              Larry Wall for Perl5 on GitHub: *utf8.h*</code></pre>\n<p>A comment reads: A note on nomenclature: The term UTF-8 is used loosely and inconsistently in Perl documentation ... perl uses an extension of UTF-8 to represent code points that Unicode considers illegal.</p>\n<p>.</p>\n<h3 id=\"tldr-examples-8\">TLDR / Examples</h3>\n<p>| ... |  |  |  |  |  |  |  |  |  |  | \n| 5 | <code>111110xx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  |  |  |  | \n| 6 | <code>1111110x</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  |  |  | \n| 7 | <code>11111110</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  |  | \n| 13 | <code>11111111</code> | <code>10000000</code> | <code>10000xxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | ... | <code>10xxxxxx</code> | </p>\n<h3 id=\"properties-8\">Properties</h3>\n<h4 id=\"bit-counts-6\">Bit Counts</h4>\n<p>There are no code units of length 8, 9, 10, 11, or 12. Nor are there any of length 14 or beyond.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>code unit length</th><th>number of content bits</th></tr></thead><tbody><tr><td><code>n = 1</code></td><td><code>7</code></td></tr><tr><td><code>1 &lt; n &lt; 8</code></td><td><code>5n+1</code></td></tr><tr><td><code>n = 13</code></td><td><code>63/64 used, up to 72?</code></td></tr></tbody></table></div>\n<h4 id=\"why-13-bytes-instead-of-12\">Why 13 Bytes Instead Of 12?</h4>\n<pre><code>              Encoding the maximum possible signed 64-bit integer, 0x7FFF_FFFF_FFFF_FFFF, in Perl utf8 returns `11111111` `10000000` `10000111` ... `10111111`. Even if the maximum possible unsigned 64-bit integer, 0xFFFF_FFFF_FFFF_FFFF, were encodable it could fit into the lower 11 bytes. So why use 13 bytes instead of 12, with the second byte always `10000000`?\n            \n\n              My guess is that the designers were going to use a UTF-Infinity style `11111111` `11111000` opening, but then they recognized that they would lose self-synchronization because the second byte looks like the start of a 5-byte utf8 sequence. Therefore they swapped the second byte to a `10000000`. They could have removed it and used 12 bytes, but that would preclude any future reunification with e.g. UTF-8000, which with 12 bytes has only `5 * 12 + 1 = 61` content bits which is less than 64, but with 13 bytes has `5 * 13 + 1 = 66` which is sufficient.</code></pre>\n<h4 id=\"information-rate-6\">Information Rate</h4>\n<p>On the surface 13-byte, 72-bit code units have an information rate of <code>72 / (13*8) = 72 / 104 = 69%</code>, which is better than that of UTF-8000 (<code>62.5%</code>).</p>\n<p>However as only 64 of those 72 bits are used in encoding 64 bit numbers, with the whole of the first continuation byte never being used, the information rate is closer to <code>64 / (13*8) = 64 / 104 = 61.5%</code>, which is worse than UTF-8000.</p>\n<h4 id=\"self-synchronization-4\">Self-Synchronization</h4>\n<pre><code>              This is effectively the same as UTF-8000. All continuation bytes have a `10` self-synchronization prefix, and the 13-byte start byte `11111111` has a `11` self-synchronization prefix.</code></pre>\n<h4 id=\"self-punctuation-4\">Self-Punctuation</h4>\n<pre><code>              The first byte of a 13-byte code unit being `11111111` characterizes it as a special case, providing self-punctuation. This is similar to ASCII being a special case with its characterizing prefix of `0` in the highest bit.</code></pre>\n<h4 id=\"not-infinitely-extensible\">Not Infinitely Extensible</h4>\n<p>Because Perl only supports up to 64-bit numbers without a specialized <code>bigint</code> module, it was sensible of them to cap their extension of UTF-8 to a finite number of bytes. It&#39;s not the prettiest however, and I&#39;m not sure why they chose 13 bytes when 12 would suffice. CPU alignment if they don&#39;t store the predictable start byte of all 1s?</p>\n<h3 id=\"verdict-7\">Verdict</h3>\n<p>Inextensible, providing only one special case beyond 7-byte UTF-8</p>\n<p> to encode 64-bit numbers, and is thus not widely known or supported.</p>\n<h2 id=\"ucs-x\">UCS-X</h2>\n<p>UCS-X proposes three extensions for each of UTF-8, UTF-16, UTF-32, for a total of nine specifications, twelve including the existing base specifications!</p>\n<p>I have so far only investigated the UTF-8 extensions, as they are all dense reads, and that is what we summarize in this section, with our main contribution being bitwise color highlighting.</p>\n<p>At a glance the UTF-16 extensions look like they break syntax with base UTF-16, whereas our UTF-16K proposal does not. Ours only semantically reinterprets the high surrogates <code>U+DB00</code> to <code>U+DB3F</code>. I will have a look at UCS-X&#39;s UTF-16 and UTF-32 extensions when I get time, to see if they contain anything interesting, or if I&#39;m wrong.</p>\n<h3 id=\"source-8\">Source</h3>\n<pre><code>              Tom Bishop and Richard Cook on ucsx.org: *The UCS-X Family of UCS Extensions (Draft Proposal)*</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>UTF-8</th><th>UTF-16</th><th>UTF-32</th></tr></thead><tbody><tr><td>UTF-G-8</td><td>UTF-G-16</td><td>UTF-G-32</td></tr><tr><td>UTF-E-8</td><td>UTF-E-16</td><td>UTF-E-32</td></tr><tr><td>UTF-∞-8</td><td>UTF-∞-16</td><td>UTF-∞-32</td></tr></tbody></table></div>\n<h3 id=\"tldr-examples-9\">TLDR / Examples</h3>\n<h4 id=\"utf-g-8\">UTF-G-8</h4>\n<p>The same as original 6-byte UTF-8 (RFC 2279) by Ken Thompson and Rob Pike, the same as UTF-8000.</p>\n<p>| ... |  |  |  |  |  |  | \n| 5 | <code>111110xx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  | \n| 6 | <code>1111110x</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | </p>\n<h4 id=\"utf-e-8\">UTF-E-8</h4>\n<p>The same as Perl utf8.</p>\n<p>| ... |  |  |  |  |  |  |  |  |  |  | \n| 7 | <code>11111110</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> |  |  |  | \n| 13 | <code>11111111</code> | <code>10000000</code> | <code>10000xxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | ... | <code>10xxxxxx</code> | </p>\n<h4 id=\"utf-8\">UTF-∞-8</h4>\n<p>This extension&#39;s code units are best characterized by the number of hex digits that the <code>U+...XXXX</code> codepoint representation consists of. The idea is that by adding on two more continuation bytes, which contain 12 content bits, one can add three more hex digits to the <code>U+...XXXX</code> codepoint representation.</p>\n<p>The 18 hex digit, 71 and 72 content bit cases are handled specially, in the transition region of extending from UTF-E-8.</p>\n<pre><code>              Otherwise, to encode an integer N: Subtract 18 from the number of hex digits in the integer&#39;s `U+...XXX` codepoint representation. Store this number in one or more low length-storage bytes</code></pre>\n<p> of the form <code>1010xxxx</code>. Precede these low length-storage bytes</p>\n<p> with the constant high length-storage bytes</p>\n<p> <code>10110100</code> (0xB4), where the number of high length-storage bytes is one less than the number of low length-storage bytes. Now precede this with the constant full start byte <code>11111111</code>. Now succeed all of this with continuation bytes that store the content bits, of the form <code>10xxxxxx</code>. These content bytes come in pairs, and the number of pairs should be one third of the number of hex digits, rounded up to the next integer if necessary.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>hex digits</th><th>content bits</th><th>bytes</th><th></th><th></th><th></th><th></th><th></th><th></th><th></th><th></th><th></th><th></th><th></th></tr></thead><tbody><tr><td>...</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>18</td><td>71</td><td>13</td><td><code>11111111</code></td><td><code>100xxxxx</code></td><td><code>10xxxxxx</code></td><td><code>10xxxxxx</code></td><td>...</td><td><code>10xxxxxx</code></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>18</td><td>72</td><td>14</td><td><code>11111111</code></td><td><code>10100000</code></td><td><code>101xxxxx</code></td><td><code>10xxxxxx</code></td><td>...</td><td><code>10xxxxxx</code></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>19</td><td>76</td><td>16</td><td><code>11111111</code></td><td><code>10100001</code></td><td><code>10000000</code></td><td><code>1000xxxx</code></td><td><code>10xxxxxx</code></td><td>...</td><td><code>10xxxxxx</code></td><td></td><td></td><td></td><td></td></tr><tr><td>20</td><td>80</td><td>16</td><td><code>11111111</code></td><td><code>10100010</code></td><td><code>100000xx</code></td><td><code>10xxxxxx</code></td><td><code>10xxxxxx</code></td><td>...</td><td><code>10xxxxxx</code></td><td></td><td></td><td></td><td></td></tr><tr><td>21</td><td>84</td><td>16</td><td><code>11111111</code></td><td><code>10100011</code></td><td><code>10xxxxxx</code></td><td><code>10xxxxxx</code></td><td><code>10xxxxxx</code></td><td>...</td><td><code>10xxxxxx</code></td><td></td><td></td><td></td><td></td></tr><tr><td>...</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>31</td><td>124</td><td>24</td><td><code>11111111</code></td><td><code>10101101</code></td><td><code>10000000</code></td><td><code>1000xxxx</code></td><td><code>10xxxxxx</code></td><td>...</td><td><code>10xxxxxx</code></td><td></td><td></td><td></td><td></td></tr><tr><td>32</td><td>128</td><td>24</td><td><code>11111111</code></td><td><code>10101110</code></td><td><code>100000xx</code></td><td><code>10xxxxxx</code></td><td><code>10xxxxxx</code></td><td>...</td><td><code>10xxxxxx</code></td><td></td><td></td><td></td><td></td></tr><tr><td>33</td><td>132</td><td>24</td><td><code>11111111</code></td><td><code>10101111</code></td><td><code>10xxxxxx</code></td><td><code>10xxxxxx</code></td><td><code>10xxxxxx</code></td><td>...</td><td><code>10xxxxxx</code></td><td></td><td></td><td></td><td></td></tr><tr><td>34</td><td>136</td><td>28</td><td><code>11111111</code></td><td><code>10110100</code></td><td><code>10100001</code></td><td><code>10100000</code></td><td><code>10000000</code></td><td><code>1000xxxx</code></td><td><code>10xxxxxx</code></td><td>...</td><td><code>10xxxxxx</code></td><td></td><td></td></tr><tr><td>35</td><td>140</td><td>28</td><td><code>11111111</code></td><td><code>10110100</code></td><td><code>10100001</code></td><td><code>10100001</code></td><td><code>100000xx</code></td><td><code>10xxxxxx</code></td><td><code>10xxxxxx</code></td><td>...</td><td><code>10xxxxxx</code></td><td></td><td></td></tr><tr><td>36</td><td>144</td><td>28</td><td><code>11111111</code></td><td><code>10110100</code></td><td><code>10100001</code></td><td><code>10100010</code></td><td><code>10xxxxxx</code></td><td><code>10xxxxxx</code></td><td><code>10xxxxxx</code></td><td>...</td><td><code>10xxxxxx</code></td><td></td><td></td></tr><tr><td>...</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>271</td><td>1084</td><td>186</td><td><code>11111111</code></td><td><code>10110100</code></td><td><code>10101111</code></td><td><code>10101101</code></td><td><code>10000000</code></td><td><code>1000xxxx</code></td><td><code>10xxxxxx</code></td><td>...</td><td><code>10xxxxxx</code></td><td></td><td></td></tr><tr><td>272</td><td>1088</td><td>186</td><td><code>11111111</code></td><td><code>10110100</code></td><td><code>10101111</code></td><td><code>10101110</code></td><td><code>100000xx</code></td><td><code>10xxxxxx</code></td><td><code>10xxxxxx</code></td><td>...</td><td><code>10xxxxxx</code></td><td></td><td></td></tr><tr><td>273</td><td>1092</td><td>186</td><td><code>11111111</code></td><td><code>10110100</code></td><td><code>10101111</code></td><td><code>10101111</code></td><td><code>10xxxxxx</code></td><td><code>10xxxxxx</code></td><td><code>10xxxxxx</code></td><td>...</td><td><code>10xxxxxx</code></td><td></td><td></td></tr><tr><td></td><td></td><td></td><td><code>11111111</code></td><td><code>10110100</code></td><td><code>10110100</code></td><td><code>1010xxxx</code></td><td><code>1010xxxx</code></td><td><code>1010xxxx</code></td><td><code>10000000</code></td><td><code>1000xxxx</code></td><td><code>10xxxxxx</code></td><td>...</td><td><code>10xxxxxx</code></td></tr><tr><td></td><td></td><td></td><td><code>11111111</code></td><td><code>10110100</code></td><td><code>10110100</code></td><td><code>1010xxxx</code></td><td><code>1010xxxx</code></td><td><code>1010xxxx</code></td><td><code>100000xx</code></td><td><code>10xxxxxx</code></td><td><code>10xxxxxx</code></td><td>...</td><td><code>10xxxxxx</code></td></tr><tr><td></td><td></td><td></td><td><code>11111111</code></td><td><code>10110100</code></td><td><code>10110100</code></td><td><code>1010xxxx</code></td><td><code>1010xxxx</code></td><td><code>1010xxxx</code></td><td><code>10xxxxxx</code></td><td><code>10xxxxxx</code></td><td><code>10xxxxxx</code></td><td>...</td><td><code>10xxxxxx</code></td></tr><tr><td>...</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr></tbody></table></div>\n<pre><code>              The maximum integer that `n` low length-storage bytes can store is `16^n - 1`, which is always congruent to `0 mod3`. Thus the number of code units in each family of `n` low length-storage bytes is `(16^n - 1) - (16^(n-1) - 1)` which is also always congruent to `0 mod3`. In the base case of `n = 1`, the maximum low length-storage byte `10101111` (33 hex digits) is succeeded by the bytes\n              `10xxxxxx`\n              `10xxxxxx`, case 3/3 of the repeating pattern of mandatory content bits placement in the highest two content bytes. Thus we conclude inductively that each family of code units with `n` low length-storage bytes ends the same way, with\n              `10101111`\n              [ `10101111` ... ]\n              `10xxxxxx`\n              `10xxxxxx`. This ensures clean transitions from `n` to `n+1` low length-storage byte families.</code></pre>\n<h3 id=\"properties-9\">Properties</h3>\n<h4 id=\"bit-counts-7\">Bit Counts</h4>\n<div class=\"table-wrap\"><table><thead><tr><th>variant</th><th>code unit length</th><th>number of content bits</th><th>same as</th></tr></thead><tbody><tr><td>ASCII</td><td><code>n = 1</code></td><td><code>7</code></td><td>UTF-8000</td></tr><tr><td>UTF-8</td><td><code>2 ≤ n ≤ 4</code></td><td><code>5n+1</code></td><td>UTF-8000</td></tr><tr><td>UTF-G-8</td><td><code>5 ≤ n ≤ 6</code></td><td><code>5n+1</code></td><td>UTF-8000</td></tr><tr><td>UTF-E-8</td><td><code>n = 7</code></td><td><code>5n+1 ( = 36)</code></td><td>UTF-8000</td></tr><tr><td>UTF-E-8</td><td><code>n = 13</code></td><td><code>63</code></td><td>Perl utf8</td></tr><tr><td>UTF-∞-8</td><td><code>n = 13</code></td><td><code>71</code></td><td></td></tr><tr><td>UTF-∞-8</td><td><code>n = 14</code></td><td><code>72</code></td><td></td></tr></tbody></table></div>\n<p>and then further for UTF-∞-8:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>number of hex digits</th><th>number of content bits</th><th>code unit length</th></tr></thead><tbody><tr><td><code>n ≥ 19</code></td><td><code>4n</code></td><td><code>2(⌊log</code>&lt;sub&gt;16&lt;/sub&gt; (n-18)⌋ + 1) + 2⌈n/3⌉</td></tr></tbody></table></div>\n<p>UTF-G-8 stores 31 bits, enough to encode positive signed 32-bit integers.</p>\n<p>UTF-E-8 stores 63 bits, enough to encode positive signed 64-bit integers.</p>\n<p>For UTF-∞-8 this is unlimited.</p>\n<h4 id=\"information-rate-7\">Information Rate</h4>\n<pre><code>              Asymptotically `4n / (8 * 2(⌊log` tends towards &lt;sub&gt;16&lt;/sub&gt;(n-18)⌋ + 1) + 2⌈n/3⌉)`3/4 = 75%`, which is better than that of UTF-8000 (`62.5%`). That&#39;s the power of the low length-storage bytes `1010xxxx` using all possible combinations of bits, whereas UTF-8000&#39;s start bits use only unary codewords. The `3/4 = 6/8` is representative of the content bytes.</code></pre>\n<h4 id=\"self-synchronization-5\">Self-Synchronization</h4>\n<pre><code>              Every byte beyond the first begins with the continuation prefix `10`, ensuring self-synchronization.</code></pre>\n<h4 id=\"self-punctuation-5\">Self-Punctuation</h4>\n<pre><code>              The high length-storage bytes `10110100` provide self-punctuation. They tell us to keep reading a stream for them until we reach a low length-storage byte.</code></pre>\n<h4 id=\"byte-map-3\">Byte Map</h4>\n<pre><code>              I was going to create one of these, but then I realized that unlike UTF-8000, UTF-∞-8 reuses bytes depending on context. For example 0xB4 can be\n              `10110100`\n              or\n              `10110100`, and 0xAX can be\n              `1010xxxx`\n              or\n              `1010xxxx`.</code></pre>\n<h4 id=\"strcmp-3-ordering-2\"><code>strcmp(3)</code> Ordering</h4>\n<pre><code>              The choice of 13 byte code units being limited to 71 bits leads to a second byte of the form\n              `100xxxxx`. This, and the choice of low length-storage bytes being of the form `1010xxxx`, high length-storage bytes being `10110100`, and the use of a unary-code-like sequence of high length-storage bytes for self-punctuation, means that UTF-∞-8 preserves `strcmp` ordering.\n            \n\n              It seems that any `1011xxxx` 0xBX byte could have been used for the high length-storage bytes, and 0xB4 before</code></pre>\n<p> is just a fun choice.</p>\n<h4 id=\"bom-support-2\">BOM Support</h4>\n<p>For the same reasons as UTF-8000, BOM support is maintained.</p>\n<h3 id=\"verdict-8\">Verdict</h3>\n<p>It works, but it&#39;s quite complicated. It took me an entire day to figure out how it works, and to calculate its stats. I much prefer the simplicity of UTF-8000.</p>\n<pre><code>              The high length-storage bytes `10110100` provide self-punctuation and `strcmp` support, but somehow feel wasteful, taking up eight bits each, and are exceptional, with none of the other 0xBX bytes being used in a similar way. That being said, asymptotically UTF-∞-8 has a better information rate than UTF-8000.</code></pre>\n<p>The mechanism of subtracting 18 from the number of hex digits and stuffing them into the low length-storage bytes reminds me a little of UTF-16, subtracting 0x10000 from the codepoint value and stuffing that into the surrogate bytes.</p>\n<h2 id=\"owl-s-corrected\">Owl&#39;s Corrected</h2>\n<p> UTF-8</p>\n<pre><code>            Owl suggests a corrected</code></pre>\n<p> version of UTF-8 with many radical changes. Many of these are opinionated, such as removing most control codes from C0. Many are technical, such as precluding the concept of overlong encodings similarly to UTF-16, by an <code>n</code>-byte code unit decoding to the binary number stored in the content bits added to the upper bound of codepoint values from <code>n-1</code>-byte code units.</p>\n<p>Even notwithstanding the established dominance of UTF-8, I still disagree with almost everything in the document. But, he does come the closest to discovering the structure of UTF-8000&#39;s code units.</p>\n<h3 id=\"source-9\">Source</h3>\n<pre><code>              Zachary Weinberg on Owl&#39;s Portfolio: *Corrected UTF-8*</code></pre>\n<h3 id=\"tldr-examples-10\">TLDR / Examples</h3>\n<p>| ... |  |  |  |  |  | \n| 6 | <code>1111110x</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | ... | <code>10xxxxxx</code> | \n| 7 | <code>11111110</code> | <code>10xxxxxx</code> | <code>10xxxxxx</code> | ... | <code>10xxxxxx</code> | \n| 8 | <code>11111111</code> | <code>110xxxxx</code> | <code>10xxxxxx</code> | ... | <code>10xxxxxx</code> | \n| 9 | <code>11111111</code> | <code>1110xxxx</code> | <code>10xxxxxx</code> | ... | <code>10xxxxxx</code> | \n| ...? |  |  |  |  |  | </p>\n<pre><code>              If Owl had specified the second-highest bit of his continuation start bytes to be a `0` instead of a `1` then he would have beat me to UTF-8000! So close, but so far.</code></pre>\n<h3 id=\"properties-10\">Properties</h3>\n<h4 id=\"no-self-synchronization-2\">No Self-Synchronization</h4>\n<pre><code>              In the 8-byte code unit example, there is no way to distinguish the second byte `110xxxxx` from the first byte of a 2-byte code unit `110xxxxx`. This generalizes beyond just 8-byte code units. This specific example could also encode `11000001` (0xC1), which UTF-8 cannot, which may trip UTF-8 compatibility stress-tests.</code></pre>\n<h4 id=\"bom-collision\">BOM Collision</h4>\n<pre><code>              Owl acknowledges that his extension may lead to issues with the UTF-16 BOM, as (presumably?) his extension permits `11111111` `11111110` (0xFF 0xFE), the little-endian UTF-16 BOM.</code></pre>\n<h3 id=\"verdict-9\">Verdict</h3>\n<p>Owl acknowledges leaving that extension for the future</p>\n<p> with respect to going beyond the 6-byte old RFC 2044 version of UTF-8, showing humility and acknowledging his design&#39;s flaws.</p>\n<p>I don&#39;t wish to dunk on his document too hard, but I&#39;m greatly relieved that he failed to derive the infinite extension mechanism. It&#39;s not just for my ego&#39;s sake, but because I do not wish for UTF-8(000) to be associated with all the other junk in his specification.</p>\n<h2 id=\"do-nothing\">Do Nothing</h2>\n<p>Why bother publishing this now and making so much noise? As of Unicode Version 17.0, September 9th 2025, only 299,448 of 1,114,112 (27%) codepoints have been designated.</p>\n<pre><code>              We choose to go to the Moon in this decade and do the other things, not because they are easy, but because they are hard, because that goal will serve to organize and measure the best of our energies and skills, because that challenge is one that we are willing to accept, one we are unwilling to postpone, and one we intend to win...</code></pre>\n<ul><li>It is a great exercise in coding theory.</li><li>Nobody else seems to have figured it out, as only worse rejected alternatives have been previously proposed.</li><li>If we wait until we run out of codepoints, one of those rejected alternatives may be hastily implemented just because it already exists. Granted, compatibility with UTF-16 would be broken, and first codepoints with five and six byte UTF-8 representations as per RFC 2044 could be satisfactory without needing UTF-8000.</li><li>The current upper bound of <code>U+10FFFF</code> on codepoints is entirely due to UTF-16&#39;s maximum capacity. UTF-16 and<code>wchar_t</code> is legacy Windows tech. If the future demands a larger set of codepoints then we should allow ourselves to not be held back.</li><li>Who doesn&#39;t like freedom, the ability to encode any integer (unsigned or signed) that we want to?</li><li>I feel in charge of the intellectual property to an extent, and as such I have copylefted it, rather than allowing a tyrant to (re-)discover it and publish it on their restrictive terms.</li><li><del>Fortune and</del> glory. I figured this out<em>myself</em> in the era of the rise of AI. I&#39;ll settle for the credit lol.</li><li>I will be submitting this work to 3b1b&#39;s Summer of Mathematics Exposition 2026.</li></ul>\n<h3 id=\"verdict-10\">Verdict</h3>\n<p>Publish.</p>\n<h2 id=\"feedback\">Feedback</h2>\n<p>Feedback is welcome, by email or on GitHub, if you have any improvements or questions. I plan on reaching out to people in phases to get the most UTF-proximal</p>\n<p> feedback first. Selected feedback may go in this section.</p>\n<h3 id=\"ken-thompson\">Ken Thompson</h3>\n<p>Ken Thompson replied to my email. That&#39;s really cool! Here is the correspondence:</p>\n<h2 id=\"emails\">emails</h2>\n<pre><code>Date: Jul 4, 2026, 1:13 AM\nFrom: Jay Berry &lt;&gt;\nTo: Ken Thompson &lt;&gt;\nSubject: I have extended UTF-8 infinitely!\nHi Ken,\nI thought you might be interested to see how (infinitely) far one can\npush UTF-8, without introducing any new special cases, and while\nmaintaining all properties like self-synchronization, `strcmp(3)`\nordering, n-byte multibyte code units having 5n+1 content bits, etc.\nI have put a one-page document on my\n[website](https://utf-8000.jb2170.com/) explaining the spec. The TLDR\nsection should be sufficient to see what&#39;s going on, splitting the\nself-synchronization bits from the self-punctuation bits, and allowing\nthe self-punctuation bits to roll over into continuation bytes.\nI have searched high and low on the internet to try to make sure that\nI have not *re*discovered this, that I am not unduly taking credit for\nit. It seems to be an original thought. I have also fairly analyzed a\nfew rejected alternatives but they all lose key properties.\nCan I ask: Did you or Rob Pike or anyone else working on FSS-UTF /\nUTF-8 intend for it to be *this* extensible / future-proof? You did a\nreally good job! At this rate it will still be around in many\ncenturies&#39; time.\nHappy Fourth of July! Consider this a 250th birthday gift from Great\nBritain (if you&#39;d not already thought of it while designing UTF-8 back\nin the 90s lol).\nThanks,\nJay Berry\n---\nDate: Jul 16, 2026, 11:29 PM\nFrom: Ken Thompson &lt;&gt;\nTo: Jay Berry &lt;&gt;\nSubject: Re: I have extended UTF-8 infinitely!\nyour first 2 extensions (5 and 6 bytes) were clearly envisioned.\nthe standard (up to 4 bytes) was created to cover the size of\nunicode. i thought any more description would be a waste of\npaper. i think your extension from 7 to 8 bytes is a little hoaky.\ni requires reading the whole string rather than &quot;knowing&quot; the\nnumber of follow on bytes. so, i think the only thing new is the\n7 byte version.\ni appreciate the mail, but i really dont think it is useful. it is\nlike replacing ipv6 with ipv50.\n---\nDate: Jul 17, 2026, 11:13 PM\nFrom: Jay Berry &lt;&gt;\nTo: Ken Thompson &lt;&gt;\nSubject: Re: I have extended UTF-8 infinitely!\nHi Ken,\nThanks for the reply!\nI agree with the &#39;ipv50&#39; remark haha. Even if we exhaust the existing\n1,112,064 possible Unicode codepoints, going back to your original\n6-byte UTF-8 proposal yields over 2 billion codepoints (31 bits),\nwhich would be sufficient for a long while, without needing\ncontinuation-start bytes.\nI&#39;m submitting UTF-8000 to the 2026 [Summer of Math\nExposition](https://some.3b1b.co/). I think it&#39;s still worth sharing\nif it inspires those interested in maths / computer science, even\nthough it may never be used in our lifetimes.\nDo you mind if I include this email chain in the feedback section? I\ndecided to first ask the creator of UTF-8 (yourself), then the authors\nof the alternatives that I&#39;ve critiqued, then the general public.\nThanks,\nJay\n---\nDate: Jul 18, 2026, 5:41 AM\nFrom: Ken Thompson &lt;&gt;\nTo: Jay Berry &lt;&gt;\nSubject: Re: I have extended UTF-8 infinitely!\nyou can use the reply.\n                </code></pre>\n<pre><code>          Only the zigzag signed variant requires reading the entire code unit (really the last bit of the last byte) to perform `strcmp` checking, not normal UTF-8000, but yes that&#39;s a good point that he&#39;s observed.</code></pre>\n<h3 id=\"rejected-alternatives-authors\">Rejected Alternatives Authors</h3>\n<p>I have emailed the authors of the rejected alternatives that I&#39;ve reviewed, to see what are their critiques of mine.</p>\n<h2 id=\"first-email\">first email</h2>\n<pre><code>Date: 2 Aug 2026, 16:14\nFrom: Jay Berry &lt;&gt;\nTo: Mashpoe          (UTF-Infinity) &lt;&gt;,\n    Larry Wall       (Perl utf8)    &lt;&gt;,\n    Tom Bishop       (UCS-X)        &lt;&gt;,\n    Richard Cook     (UCS-X)        &lt;&gt;,\n    Zachary Weinberg (Owl)          &lt;&gt;\nSubject: Unlimited UTF-8 | UTF-8000\nHi everyone!\nI believe that I have discovered the &quot;correct&quot; way to extend UTF-8 infinitely,\nwithout introducing any new special cases, and while maintaining all properties\nlike self-synchronization, self-punctuation, `strcmp(3)` ordering, n-byte multibyte\nunits having 5n+1 content bits, etc. I&#39;ve codenamed it &quot;UTF-8000&quot; or &quot;UTF-8K&quot;.\nI have put a one-page document on my [website](https://utf-8000.jb2170.com/)\nexplaining the spec. The TLDR section should be sufficient to see what&#39;s going on,\nsplitting the self-synchronization bits from the self-punctuation bits, and\nallowing the self-punctuation bits to roll over into continuation bytes.\nReference implementation in Python is available on\n[GitHub](https://github.com/UTF-8000/UTF-8000-Python) which can be installed\nwith `$ pipx install UTF-8000`.\nI noticed that each of you have attempted to extend UTF-8 in different ways,\nand I have constructively reviewed each of them in the\n[rejected alternatives](https://utf-8000.jb2170.com/#sec-rejected-alternatives)\nsection of my spec. I thought you might be interested / maybe you have some\nfeedback for mine.\n- [UTF-Infinity](https://utf-8000.jb2170.com/#sec-rejected-utf-infinity)   by Mashpoe\n- [Perl utf8](https://utf-8000.jb2170.com/#sec-rejected-perl-utf8)         by Larry Wall\n- [UCS-X](https://utf-8000.jb2170.com/#sec-rejected-ucs-x)                 by Tom Bishop and Richard Cook\n- [Owl&#39;s &quot;Corrected&quot; UTF-8](https://utf-8000.jb2170.com/#sec-rejected-owl) by Zachary Weinberg\nI emailed Ken Thompson, creator of UTF-8 (and Unix!) to see what he thinks,\nand I got a reply! The exchange is on the\n[website](https://utf-8000.jb2170.com/#sec-feedback-ken-thompson). His remark\n&quot;it is like replacing ipv6 with ipv50&quot; is funny to me and should set a\nnot-too-serious atmosphere for this whole discussion.\nNonetheless I still think that UTF-8000 is fun and educational anyways, an exercise\nin coding theory, and I&#39;ll be submitting it to 3b1b&#39;s 2026\n[Summer of Math Exposition](https://some.3b1b.co/). But first I thought that it\nwould be proper to email the people whose work I&#39;ve reviewed.\nThanks,\nJay Berry\n                </code></pre>\n<pre><code>          #### Zachary Weinberg (Owl)</code></pre>\n<h2 id=\"emails-2\">emails</h2>\n<pre><code>Date: 2 Aug 2026, 19:18\nFrom: Zachary Weinberg &lt;&gt;\nTo: Jay Berry &lt;&gt;\nSubject: Re: Unlimited UTF-8 | UTF-8000\nOn Sun, Aug 2, 2026, at 11:14 AM, Jay Berry wrote:\n&gt; I believe that I have discovered the &quot;correct&quot; way to extend UTF-8\n&gt; infinitely, without introducing any new special cases, and while\n&gt; maintaining all properties like self-synchronization, self-\n&gt; punctuation, `strcmp(3)` ordering, n-byte multibyte units having 5n+1\n&gt; content bits, etc. I&#39;ve codenamed it &quot;UTF-8000&quot; or &quot;UTF-8K&quot;.\nHey, thanks for reaching out.  I&#39;m delighted to see that I am not the\nonly one fed up with the artificial limitation of UTF-8&#39;s encoding space\nto match UTF-16. I may actually revise my proposal to adopt your trick\nfor preserving self-synchronization even when the start bits extend\npast the end of the first byte.\nI think you&#39;re not taking the value of *eliminating* overlength\nencodings seriously enough, though.  Yeah, that&#39;s the most complicated\npart of my proposal, and the part that means Corrected UTF-8 doesn&#39;t\ncorrespond to IETF UTF-8 for anything but the ASCII page, but it&#39;s also\nthe part that means Corrected UTF-8 decoders *cannot* be a vehicle for\npath-smuggling attacks on network services, and therefore I consider it\nsecond in importance only to lifting the artificial plane limit.\nMaking it impossible to encode surrogates is also important for\nsecurity reasons; the only way I could be persuaded to not do that is\nif there was any chance that the surrogates might get *reassigned* as\nordinary characters in a couple decades, once UTF-16 is truly dead and\nburied ... and I think the odds of that ever happening are far lower\nthan the odds of the Unicode Consortium backing down on their &quot;never will\nthere be more than 17 planes&quot; policy.\nI&#39;ve mostly come around to agree with you on the C1 controls, though.\nOmitting them doesn&#39;t really help anything.  At the time I wrote the\noriginal document (some years before I posted it on my website) I was\nstill working at a browser company and mislabeled or mistranscoded\nWindows-1252 was a regular headache; but I have the impression that\nhas become much less common over the past decade and a half, and\n&quot;this can represent any Unicode codepoint with an official non-surrogate\nassignment&quot; *is* a desirable property for anything calling itself an UTF.\nzw\n---\nDate: 3 Aug 2026, 23:34\nFrom: Jay Berry &lt;&gt;\nTo: Zachary Weinberg &lt;&gt;\nSubject: Re: Unlimited UTF-8 | UTF-8000\nHi Zack,\nThanks for the reply!\n&gt; I may actually revise my proposal to adopt your trick\n&gt; for preserving self-synchronization even when the start bits extend\n&gt; past the end of the first byte.\nSelf-synchronization was indeed a main feature that Ken Thompson figured out\nin fixing FSS-UTF:</code></pre>\n<p>0vvvvvvv\n10vvvvvv 1vvvvvvv\n110vvvvv 1vvvvvvv 1vvvvvvv\n...</p>\n<pre><code>(in which one couldn&#39;t tell the difference between eg a 2-byte start byte\n`10|vvvvvv` and a continuation byte `1|0vvvvvv`) to UTF-8:</code></pre>\n<p>0vvvvvvv\n110vvvvv 10vvvvvv\n1110vvvv 10vvvvvv 10vvvvvv\n...</p>\n<pre><code>in which the self-synchronization prefixes `0`, `10`, and `11` are distinct.\n[History of FSS-UTF -&gt; UTF-8](https://www.cl.cam.ac.uk/~mgk25/ucs/utf-8-history.txt\n#:~:text=10zzzzzz%201yyyyyyy). It is definitely well worth keeping :)\n&gt; I think you&#39;re not taking the value of *eliminating* overlength\n&gt; encodings seriously enough, though.\nHaving n-byte (modified) UTF-8 decode to (value of content bits) plus\n(1 more than maximum that (n-1)-byte UTF-8 can encode) i.e. the offsets in\nyour specification, sounds good at first, but if n is big then the offset accrues:\n2^7 + 2^11 + 2^16 + ... + 2^(5(n-1)+1). We can write this as\n2^7 + 2^11 * ((2^5)^0 + (2^5)^1 + ... + (2^5)^(n-3)) as a geometric series and\nexplicitly compute it as 2^7 + 2^11 * (32^(n-2) - 1) / 31 for n &gt;= 3,\nbut that division is a bit &#39;icky&#39; compared to addition subtraction multiplication\nand bitshifting.\n```py\ndef offset(n: int) -&gt; int:\n    if n == 1:\n        return 0\n    elif n == 2:\n        return (1 &lt;&lt; 7)\n    else:\n        return (1 &lt;&lt; 7) + (1 &lt;&lt; 11) * ((1 &lt;&lt; (5 * (n - 2))) // 31)</code></pre>\n<p>It seems a lot easier to say &quot;n-byte multibyte UTF-8 can store up to\n(5n+1)-bit codepoints&quot;, a nice instance being 3-byte UTF-8 storing 16 bits,\n1 2 and 3 byte UTF-8 exactly covering\n<a href=\"https://en.wikipedia.org/wiki/Plane_(Unicode)\" rel=\"nofollow ugc noopener\">Plane 0</a> of Unicode.\n<a href=\"https://en.wikipedia.org/wiki/UTF-1\" rel=\"nofollow ugc noopener\">UTF-1</a> was an earlier encoding that\nKen Thompson and Rob Pike tried out\n(<a href=\"https://www.youtube.com/watch?v=OmVHkL0IWk4&amp;t=14275s\" rel=\"nofollow ugc noopener\">interview</a>). It used\n<code>mod 190</code> and divisions which they disliked, and eventually FSS-UTF and UTF-8\ncame around which use simple bitwise operations to check against overlong encodings,\nand to extract the content bits. I think the anti-overlong checking is not too\ncomplicated, 2-byte UTF-8 being the only odd one out.\nUnicode offered a way forward from the ISO 8859-{1..16} diaspora of 8-bit codepages.\nUTF-8 offered ASCII forwards compatibility with better efficiency than UTF-16 using\nbyte-precision rather than word-precision. I don&#39;t think that your offset-based\nUTF-8 offers a significant upgrade. It&#39;s really just to protect noob software\ndevelopers who might write broken decoders for UTF-8 that don&#39;t do anti-overlong\nchecking and surrogate range checking.</p>\n<blockquote><p>any chance that the surrogates might get <em>reassigned</em> as\nordinary characters in a couple decades\nMy guess is that regardless of whether UTF-16 continues to live on, the\nsurrogate range will remain unencodable, due to pre-established UTF-8 parsers\nrejecting them. Though I could be wrong, as iirc IP addresses that ended in <code>.0</code>\nwere originally not allowed, but now are. It does make one wonder what those\n2048 codepoints could be assigned to...\nI&#39;ve mostly come around to agree with you on the C1 controls, though.\nThat&#39;s great. I don&#39;t think I&#39;ve ever actively used them, but yeah C1 should be\nin / remain in Unicode as a way to refer to it using codepoints. It&#39;s a bit of a\nshame that they don&#39;t (yet) have a &#39;control pictures&#39; block like\n<a href=\"https://www.compart.com/en/unicode/block/U+2400\" rel=\"nofollow ugc noopener\">C0 Does</a>.\na regular headache\nOn a similar note there&#39;s still some software that is fiddly with them. I&#39;ve got\na <a href=\"https://github.com/jb2170/better-adb-sync/issues/42\" rel=\"nofollow ugc noopener\">bug to fix</a> that I&#39;ve\nfigured out the solution to whilst messing with UTF-8(000) and encodings.\nAndroid&#39;s Toybox&#39;s <code>ls</code> outputs U+0080 to U+009F &#39;C1 control codes&#39; and\nU+00A0 &#39;non-breaking space&#39; differently depending on whether <code>ls</code> is running over\n<code>adb</code> interactively or not: interactively U+00A0 prints as <code>\\240</code> octal-escape style,\nbut non-interactively it prints as the byte <code>a0</code>, which is either a coincidentally\ndecapitated UTF-8 unit <code>c2 a0</code>, or the raw ISO-8859-1 8-bit byte. The fix is to\nuse the <code>-b</code> flag on <code>ls</code> to force an escaped style. So tldr I understand the\nfiddly-ness lol.\nThanks,\nJay</p></blockquote>\n<pre><code>              #### Tom Bishop (UCS-X)\n\n## emails\n</code></pre>\n<p>Date: 2 Aug 2026, 19:41\nFrom: Thomas Eugene Bishop &lt;&gt;\nTo: All &lt;&gt;\nSubject: Re: Unlimited UTF-8 | UTF-8000\nHi Jay,\nThanks for letting me know about your work, and the others you reference. It&#39;s\ngood to know that others are interested in extending the range of encoding.\nI&#39;ll study the proposals more when I have time. Based on first impressions, I\nhave these comments.\nAbout &quot;the &#39;correct&#39; way&quot;: maybe you mean that ironically and recognize there&#39;s\nmore than one way to do it, with trade-offs. On the other hand, you wrote,\n&quot;Nobody else seems to have figured it out, as only worse rejected alternatives\nhave been previously proposed.&quot; That sounds like an unwarranted claim that\nyou&#39;ve solved a problem nobody else was able to solve. I wish you wouldn&#39;t use\nthe word &quot;rejected&quot; to describe alternatives, since it might be misconstrued\n(maybe through an AI search) as implying a decision by an organization with some\ncapacity to accept or reject proposals. I think what you mean is that you\npersonally prefer your own proposal.\nNow that multiple solutions exist, there&#39;s room to compare them by various\ncriteria such as efficiency, simplicity, and robustness.\nYou described Larry Wall&#39;s utf8 as &quot;inextensible&quot;; that&#39;s wrong, as proved by\nits extension to UTF-∞-8. Or else, &quot;inextensible&quot; doesn&#39;t mean what I think it\nmeans. Also, my understanding is that the contrast between &quot;utf8&quot; and &quot;UTF-8&quot;\nwas intentional.\nYou wrote, &quot;UCS-X proposes three extensions for each of UTF-8, UTF-16, UTF-32,\nfor a total of nine specifications, twelve including the existing base\nspecifications!&quot; and &quot;it&#39;s quite complicated&quot;. I think this reference to 12\nspecs is an unfair criticism. The existence of multiple specs doesn&#39;t imply\ncomplexity of the encodings themselves. The complication of the existing 3 base\nspecs is obviously beyond anybody&#39;s control at this point. Of the remaining 9,\nyou can ignore 6 if you want, since they are merely simplifications of the last\n3; that is, the specs with max U+7FFFFFFF and U+7FFFFFFFFFFFFFFF are just\nsubsets of the specs with max infinity. We separated them out to support\nimplementers who might have good reasons not to go straight to infinity.\nYou wrote, &quot;At a glance the UTF-16 extensions look like they break syntax with\nbase UTF-16, ...&quot;. I don&#39;t know what you mean by &quot;break syntax&quot;, but UTF-∞-16 is\na compatible extension of UTF-16 in the sense that our spec defines &quot;compatible\nextension&quot;. It would be more responsible to postpone publishing a &quot;break syntax&quot;\nassertion until you&#39;re certain and ready to explain what you mean by it.\nYou wrote, &quot;... UTF-8000 code units can be arbitrarily large&quot; -- I think you\nmean UTF-8000 codes can be arbitrarily large. A UTF-8000 code unit is always 8\nbits, right?\nTo me, while the technical details of encoding are interesting, what&#39;s more\ninteresting is how people might eventually use extended encodings, such as to\ndefine their own characters and use them for public communication, without\nhaving to wait for official approval of each character.\nBest wishes,\nTom</p>\n<hr />\n<p>Date: 3 Aug 2026, 16:57\nFrom: Thomas Eugene Bishop &lt;&gt;\nTo: All &lt;&gt;\nSubject: Re: Unlimited UTF-8 | UTF-8000\nHi Jay,\nI wrote a script to compare the lengths of UTF-8000 and UTF-∞-8 codes, and also\ntheir &quot;start&quot; bytes. That script isn&#39;t thoroughly tested and it might be only\napproximate especially in some edge cases. With that disclaimer, it seems that\nif a USV has 47 or more digits, then UTF-8000 is longer than UTF-∞-8. If a USV\nhas 18 or more digits, the number of &quot;start bytes&quot; (needed to determine the\nlength of an entire code) is longer for UTF-8000 than for UTF-∞-8. For a USV\nwith 128 digits, UTF-8000 has 102 total bytes and 17 start bytes, while UTF-∞-8\nhas 90 total bytes and 4 start bytes. UTF-8000 does have shorter codes in some\nranges, such as for USV with 10-15 digits.\nNeither solution is optimal in terms of storage size. There are trade-offs such\nas speed of execution, simplicity, robustness, etc.\nThe number of start bytes might be important in situations where text is read\ninto a fixed-size buffer and a buffer might contain a partial code. Software\nshould be able to determine the length of an entire code by scanning a\nrelatively small number of start bytes, both for efficiency and to avoid bugs in\ncases where one code might span many buffers. This is an example of\n&quot;robustness&quot;. Another example is that protocols should enable processes to\nindicate max supported USV.\nThe term &quot;code unit&quot; has a standard definition\n(<a href=\"https://unicode.org/glossary/#code_unit\" rel=\"nofollow ugc noopener\">https://unicode.org/glossary/#code_unit</a>) that differs from yours\n(<a href=\"https://utf-8000.jb2170.com/#def-code-unit\" rel=\"nofollow ugc noopener\">https://utf-8000.jb2170.com/#def-code-unit</a>). I recommend following the standard\nto avoid confusion.\nIt&#39;s wonderful that you might bring up this topic at the Summer of Math Exposition!\nCheers,\nTom</p>\n<hr />\n<p>Date: 8 Aug 2026, 21:26\nFrom: Jay Berry &lt;&gt;\nTo: Thomas Eugene Bishop &lt;&gt;\nSubject: Re: Unlimited UTF-8 | UTF-8000\nHi Tom,\nThanks for the feedback!</p>\n<blockquote><p>About &quot;the &#39;correct&#39; way&quot;: maybe you mean that ironically and recognize\nthere&#39;s more than one way to do it, with trade-offs.\nThere are indeed other solutions such as UTF-∞-8 which preserve all properties\nlike self-synchronization, self-punctuation, strcmp order etc. However the\nreason I&#39;ve referred to it as the &#39;correct&#39; way is because in my opinion it\nlooks like the &#39;natural&#39; way to extend UTF-8, as I put in the [properties]\nsection addressing the fact that the anti-overlong mechanism works the same as\nUTF-8, with no new special cases. I find it very simple to explain (in\nretrospect) to begin with bytes endowed with self-synchronization prefixes &#39;11&#39;\nand &#39;10&#39;, and to stripe the self-punctuation bits across them.\nyou wrote, &quot;Nobody else seems to have figured it out, as only worse rejected\nalternatives have been previously proposed.&quot;\nBy that I mean that nobody else online has suggested the exact layout that\nUTF-8000 proposes, which I feel is the &#39;natural&#39; / &#39;correct&#39; one, formally\nidentifying the self-synchronization and self-punctuation bits and how to use\nthem. The verdicts on the other proposals summarize their flaws, all but UTF-∞-8\nlosing key properties of interest.\nI wish you wouldn&#39;t use the word &quot;rejected&quot; ... implying a decision by an\norganization with some capacity to accept or reject proposals\nI styled my document a bit like a [Python PEP], in which often the alternatives\nhave to be firmly disproven. That being said, yes I don&#39;t think I&#39;ve made it\nclear that this is a <em>proposal</em>, not an existing standard. In the Python\nreference implementation [readme] I added the line &quot;UTF-8000 is in no way\nendorsed by or representative of the Unicode Consortium. This is a standalone\nproject.&quot;. I think I&#39;ll copy that to the header of the website, thanks!\nYou described Larry Wall&#39;s utf8 as &quot;inextensible&quot;\nI know it looks like I&#39;m contradicting myself 10 seconds later by pointing out\nthat UCS-X extends from utf8, but what I meant is that Perl utf8 <em>on its own</em> is\ndesigned only to go up to 2^63-1. It uses the <code>FF</code> byte to start its 13-byte\nunits and doesn&#39;t specify how one could continue onwards. I also don&#39;t feel the\nneed for UTF-8000 to extend utf8 like UCS-X does, as utf8 is not used outside\nPerl, and we have the opportunity to make UTF-8000 more flexible allowing\n8,9,10,11,12 byte units with &#39;correct&#39; self-punctuation syntax (whereas utf8&#39;s\nsecond byte is just a plain 0x80).\nAlso, my understanding is that the contrast between &quot;utf8&quot; and &quot;UTF-8&quot; was intentional.\nYeah I&#39;ll remove that line about &quot;utf8&quot; vs &quot;UTF-8&quot;, thanks. I know that the\nUnicode Consortium is pedantic with referring to &#39;UTF-8&#39; using a hyphen, and I\noriginally thought that Perl was just being a bit loose with the naming. It is\nmore likely that &#39;utf8&#39; was chosen to show that it&#39;s not <em>exactly</em> &#39;UTF-8&#39;, like\nI&#39;m using &#39;UTF-8000&#39; as a codename for my proposal.\nI think this reference to 12 specs is an unfair criticism.  The existence of\nmultiple specs doesn&#39;t imply complexity of the encodings themselves.  We\nseparated them out to support implementers who might have good reasons not to\ngo straight to infinity.\nWe can group UTF-8 and UTF-G-8 together since they both follow the same style,\nand 5/6-byte UTF-8 was envisioned by Ken Thompson. As for the UTF-E-8 and\nUTF-∞-8 specifications, they are very different.\nI think that UTF-8000, which is just one specification, within which there are\n&#39;natural ranges&#39; (ie limiting to n-byte units) is a better approach. My original\nspecification for UTF-16K was going to use just one Unicode Plane, to provide\ndecent efficiency but without being too greedy in needing to claim existing\nUnicode codepoints. But then I realised that this would provide 14k+1 content\nbits, whereas if we used two planes instead of one, this would be 15k+1 content\nbits, which overlaps nicely with 5n+1 provided by UTF-8000. So I have taken some\nthought and care as to create &#39;ranges&#39; like your &#39;Giga&#39;, &#39;Exa&#39;, &#39;Inf&#39; ideas,\nwith which UTF-8 and UTF-16 can be expanded in parallel. I put this in the\n[UTF-16K spec]. I think it&#39;s a lot easier to say &quot;this is what n-byte UTF-8 and\nk-surrogate-pair UTF-16 looks like. restrict to n=3k and you have ranges that\nencode the same codepoints&quot; than to have a patchwork of different standards\nbased on what range a codepoint is in, like UCS-X eg includes Perl utf8 as\nUTF-E-8.\nI don&#39;t know what you mean by &quot;break syntax&quot;, but UTF-∞-16 is a compatible extension\nBy &#39;compatible&#39; I&#39;m thinking along the lines of backwards compatibility &quot;will\nthis throw an error in a UTF-8 / UTF-16 parser?&quot; and &quot;are we maintaining the\npre-established syntax?&quot;.\nFor UTF-8, technically one could argue that UTF-8000 and UTF-∞-8 &quot;break syntax&quot;\nby eg using the byte &#39;FF&#39;, which when fed into a UTF-8 parser will cause an\nexception. However on the other hand the byte &#39;FF&#39; causing an exception is only\ndue to the restriction to U+10FFFF on the range of codepoints for UTF-8,\nprovided one&#39;s UTF-8 extension uses the byte &#39;FF&#39;. So yes I&#39;m being a bit\nhypocritical, but I feel fine with that because bytes F{5..F} are currently\nunused by UTF-8, and the proposed syntax of UTF-8000 is predictably the same as\nUTF-8, eg wrt self-synchronization prefixes for non-ASCII first bytes being\n&#39;11&#39;, and for continuation bytes being &#39;10&#39;.\nFor UTF-16, every 16-bit word has already been used. Instead of changing the\nsyntax to use eg 1 high surrogate and (n-1) low surrogates, or like UTF-G-16 use\nn low surrogates, I decided to use a &quot;semantic reinterpretation&quot; layer on top of\nUTF-16, ASCVI-on-UTF-16. Ie, just as UTF-16 is a semantic reinterpretation of\nUCS-2, interpreting codepoints in the ranges U+D800 to U+DBFF and U+DC00 to\nU+DFFF no longer as those individual codepoint values, but rather as parts of\nsurrogate pairs, so too I decided to implement UTF-16K as a semantic\nreinterpretation of plane 9 and 10 surrogate pairs. The nice thing about this is\nthat a decoder which only understands UTF-16 can open a UTF-16K encoded file,\njust as a UCS-2 decoder can open UTF-16 files. Plane 9 and 10 surrogate pairs\nwould be displayed as UTF-16 codepoints rather than as one UTF-16K codepoint,\njust as a UCS-2 parser would show two surrogate codepoints instead of one UTF-16\ncodepoint; semantic errors rather than syntax errors. Contrast that with\nUTF-G-16, U+110000 encoded as &#39;DC04 DE80 DE00&#39;, with which the opening word may\nimmediately raise an exception in a UTF-16 parser.\nFor UTF-G-16, for ill-formed units, I am able to generate context-dependent\nerror handling behavior which leads to errors being decoded as though they are\ncorrect. I am able to cause a contradiction in your UTF-G-16 [decoding rules]:\nmake &#39;DC04&#39; both preceded by D800 (to make it trailing) and succeeded by DE80\n(to make it leading). If we were to seek to the point &#39;X&#39; in a stream &#39;X D800 Y\nDC04 DE80 DE00&#39; we would decode this as &#39;U+10004 (D800 DC04) U+FFFD (replace\nDE80) U+FFFD (replace DE00)&#39;. If we were to seek to the point &#39;Y&#39; we would\ndecode this as &#39;U+110000 (DC04 DE80 DE00)&#39;, using the low surrogate &#39;DC04&#39; and\nleaving the high surrogate &#39;D800&#39; before the seek point Y. This looks like bad\nbehavior. In UTF-8 and UTF-8000 because the first-byte and continuation-byte\nself-synchronization prefixes make their byte ranges disjoint, I don&#39;t think a\nsituation like this can happen there. Ie never will a &#39;well formed unit X\nfollowed by errors&#39; be incorrectly decoded as a &#39;well formed unit Y with perhaps\nsome junk before it&#39; if one seeks to the middle of the well formed unit &#39;X&#39;. So\ntoo UTF-16K keeps the {high surrogate | low surrogate} and {first surrogate pair\n(plane 9) | continuation surrogate pair (plane 10)} ranges disjoint which avoids\nthis issue and maintains self-synchronization at the word-level. UTF-G-16\nmuddies the water by &#39;DC04&#39; being trailing (UTF-16 surrogate pair) or leading\n(UTF-G-16 leading) dependent on previous words. This is also why a UTF-8 /\nUTF-8000 parser only ever needs to seek <em>forwards</em> to the next first byte if it\nencounters an error.\n<del>For UTF-G-16, for well formed units, something still doesn&#39;t feel right that\none might need to look backwards to determine whether eg &#39;DC04&#39; is trailing or\nleading. We do not always have backwards seeking, like on a pipe or socket, or\nat least we don&#39;t want to do backtracking like complicated regexes sometimes\ndo.</del> In well formed units we know exactly one of those conditions will be true\nand we can look forwards rather than back, right? This seems like minutiae\ncompared to the behaviour in the previous paragraph.\nBack to UTF-8, this conversation has made me realize that one could implement\n<em>private-use extensions</em> on top of Unicode / UTF-8 using ASCVI-on-UTF-8, in a\nsimilar way to UTF-16K using ASCVI-on-UTF-16. We can achieve an ASCVI-like code\nin as little as 3 bits, 8 codepoints:\n0: 000, 1: 001, 2: 010 110, 3: 010 111, 4: 011 101 100, 5: 011 101 101,\n6: 011 101 110, 7: 011 101 111, 8: 011 110 110 100, 9: 011 110 110 101, ...\nthough using more bits will of course lead to more efficient codes. The\nadvantage of this style is that it&#39;s just a semantic reinterpretation layer on\ntop of UTF-8, and will pass right through a UTF-8 parser okay. A good range of\ncodepoints to use may be some of the U+E000 to U+F8FF Plane 0 private-use\ncodepoints. This seems like a great way in which one could create their own\nautonomous set of &#39;MyUnicode&#39; codepoints M+...XXXX starting at M+0000,\nMyUnicode-on-Unicode style (as opposed to UTF-16K which uses the <em>public</em>\nUnicode range and postulates starting at U+110000). This would answer your\npoint:\nwhat&#39;s more interesting is how people might eventually use extended encodings,\nsuch as to define their own characters and use them for public communication,\nwithout having to wait for official approval of each character\nIt does somewhat go against the spirit of &quot;Uni&quot;code, which is the one-and-only\n&#39;flat&#39; layer of codepoints, to use an ASCVI layer on top of Unicode / UTF-8. One\ncan also imagine ASCVI-on-(ASCVI-on-UTF-8) if the M+...XXXX codepoints had their\n<em>own</em> private-use area which allowed further sub-encoding. It&#39;s a fun thought to\nthink of trees of Unicode embedded recursively as layers on top of each other,\nbut it would surely be a bit anarchic and low-efficiency. Therefore my main\nfocus with UTF-8000 and UTF-16K has been on how <em>Unicode</em> could expand in the\nlong run. The private-use extensions do sound fun, but may be a bit clunky when\ndecoded in a programming language, being interspersed in &#39;normal&#39; Unicode\nstrings.\nI wrote a script to compare the lengths of UTF-8000 and UTF-∞-8 codes, and\nalso their &quot;start&quot; bytes.\nYes UTF-∞-8 has shorter units in the long run, an efficiency tending towards 6/8\nwhereas UTF-8000&#39;s efficiency tends towards 5/8. It&#39;s probably easiest to point\nto UTF-8000 using a linear number of self-punctuation bits (n-1) -&gt; O(n),\nwhereas UTF-∞-8 is roughly logarithmic O(log_2(n)).\nSoftware should be able to determine the length of an entire code by scanning\na relatively small number of start bytes, both for efficiency and to avoid\nbugs in cases where one code might span many buffers.\nThis is a good point, and UTF-∞-8 is more succinct with respect to\nself-punctuation. For 33 hex-digit codepoints, UTF-∞-8 uses 2 bytes, whereas\nUTF-8000 uses 5, for 273 hex-digits UTF-∞-8 uses 4 bytes, whereas UTF-8000 uses\n37! Mogs me.\nand finally\nThe term &quot;code unit&quot; has a standard definition that differs from yours. I\nrecommend following the standard to avoid confusion.\nI realised this half way through writing the UTF-8000 spec and I&#39;m struggling to\nthink of an alternative name. I opened a GitHub [issue] to remind me to rename\nit. 😅\nSo to conclude so far:</p></blockquote>\n<ul><li>I still think that UTF-8000 is simpler to explain and more predictable than UTF-∞-8</li><li><p>UTF-∞-8 is asymptotically more efficient than UTF-8000 and requires less start</p><p>bytes (self-punctuation bytes)</p></li><li><p>UTF-G-16 (and beyond?) looks broken to me, though I haven&#39;t properly anatomized</p><p>the UTF-X-16 family of UCS-X proposals like I have for the UTF-X-8 family.</p></li><li><p>ASCVI-on-private-use-UTF-8 sounds like an okay idea for private-use extensions</p><p>if they require a large amount of &#39;codepoints&#39; (sub-encoded virtual\nmy-codepoints M+...XXXX)</p></li><li><p>I have a few remarks to change on my proposal</p><p>This has been a fun project! Thanks for the emails,\nJay</p></li></ul>\n<pre><code>              I have updated the UTF-16K specification to mention UCS-X&#39;s UTF-G-16&#39;s flawed error handling behavior.\n\n### SoME 2026\n\nI am submitting this work to 3b1b&#39;s Summer of Mathematics Exposition 2026. I hope that it is useful to some people who view it, that it is educational about coding theory, and that maybe there&#39;ll be some feedback.\n\n### General\n\nI&#39;ll put a link to this page on r/Unicode. There&#39;s lots of show-and-tell on there.\n\n### Unicode Consortium\n\nI *might* send this to the Unicode Consortium if there is good consensus from the feedback above.\n\nBut as Ken put it we don&#39;t really need IPv50 or unlimited UTF-8 right now, so I don&#39;t want to pester the Unicode Consortium when they&#39;re busy doing actually important jobs like documenting scripts, assigning codepoints, helping internationalization, etc.\n\nMaybe this specification can sit in a 250-year time capsule in the Unicode Consortium Archives for when the time is right to expand...\n\n## Reference Implementation and Tools\n\n### UTF-8000\n\nWorking reference implementation in Python with comprehensive code documentation is available on GitHub as UTF-8000/UTF-8000-Python. It can be installed as a PyPI package using `$ pipx install UTF-8000` which provides the command line utility `utf-8000(1)`.\n\nThe `$ utf-8000 info` subcommand displays info about a codepoint encoded in UTF-8000, with useful bit highlighting.\n\nThe `$ utf-8000 encode` subcommand reads codepoints from stdin and writes the raw UTF-8000 bytes to stdout.\n\nThe `$ utf-8000 decode` subcommand reads UTF-8000 bytes from stdin and feeds them to an incremental decoder, writing the decoded codepoints to stdout.\n\n### UTF-16K\n\nWorking reference implementation for UTF-16K is also available on GitHub as UTF-8000/UTF-16K-Python. It can be installed using `$ pipx install UTF-16K` which provides `utf-16k(1)` with the same subcommands as `utf-8000(1)`.\n\n## Naming\n\nIn the development phase of this project I have been using the codename UTF-8000\n\n, but I find myself increasingly drawn to UTF-8K\n\n.\n\nBelow is a comparison of different potential names, and I am open to suggestions.\n\n### UTF-8000\n\nInspired by Python 3&#39;s development codenames in PEP 3000.\n\n#### Pros\n\n- The thousand inUTF eight thousand sounds big and futuristic. This encoding scheme should also last forever!\n\n#### Cons\n\n- The zeros are repetitive.\n- Having to remember exactly three zeros to make 8000 . Some people may read it as eight hundred, or eighty thousand etc.\n- The inspiration logic doesn&#39;t exactly match up with Python because Python was moving from Python 2 to Python 3, not Python 3 to Python 3000, whereas we&#39;re going from UTF-8 to UTF-8000.\n- UTF-8 is a prefix ofUTF-8000 . Existing software which parses a string representing the encoding&#39;s name to determine the encoding may do something like`if encoding_name[:5] == &quot;UTF-8&quot;` and incorrectly short-circuit. It is for this reason that Microsoft skipped from Windows 8 to Windows 10, not creating Windows 9, because existing software may check forWindows 9 to test forWindows 95 orWindows 98 .\n\n### UTF-8K\n\nInspired by another of Python 3&#39;s development codenames Py3K\n\n / Py3k\n\n, we could use UTF-8K, the K\n\n capitalized like the UTF\n\n. It&#39;s also shorter than 8000\n\n.\n\n#### Pros\n\n- Short; just one more letter than UTF-8 .\n\n#### Cons\n\n- UTF-8 is a prefix ofUTF-8K . See here.\n\n### UTF-8 &amp; Knuckles\n\nThe K\n\n in UTF-8K\n\n reminds me of Sonic 3 &amp; Knuckles\n\n, sometimes abbreviated to S3K\n\n.\n\n                The Sonic &amp; Knuckles\n\n cartridge uses lock-on technology\n\n to extend Sonic 3\n\n and Sonic &amp; Knuckles\n\n into Sonic 3 &amp; Knuckles\n\n. In a similar way UTF-8 extends from (locks on to) ASCII, and UTF-8000 continues this extension.\n              \n\n#### Pros\n\n- Knuckles is cool.\n\n#### Cons\n\n- SEGA might not be happy, though they do seem nicer than Nintendo with respect to fanart.\n- The hardest work of extending from ASCII to UTF-8 has already been achieved by Ken Thompson and Rob Pike. Is UTF-8000 lock-on if it doesn&#39;t introduce any new special cases? Not really.\n- UTF-8 is a prefix ofUTF-8 &amp; Knuckles . See here.\n- This is just a bit of a joke-y name.\n\n### VTF-8\n\nShort for Variable Transformation Format 8\n\n.\n\n#### Pros\n\n- Removes the association with Unicode as VTF-8 is just a method of storing unsigned integers.\n- UTF-8 is not a prefix ofVTF-8 ; in fact they don&#39;t even begin with the same letter.\n- The V looks Romanesque and fits with the logo.\n\n#### Cons\n\n- The V inVTF-8 looks very similar to theU in UTF-8. At a glance, and depending on font rendering, one may not notice the difference.\n\n### STF-8 for the Signed Variant\n\nThe U\n\n in UTF-8\n\n could also be read as Unsigned\n\n, ie Unsigned Transformation Format 8\n\n. Thus we could take inspiration and write STF-8\n\n for Signed Transformation Format 8\n\n.\n\n## Logo\n\nUTF-8 does not have a logo. I had a bit of fun designing a logo for UTF-8000. I made it Romanesque but simple.\n\nIt&#39;s an eagle with UTF\n\n written across its wings and chest. The Roman numerals for 8\n\n, VIII, flank its head.\n\n                Below its claws it holds a fasces, wrapped with a continuation byte of the form\n                `10xxxxxx`. This represents the united strength of a bundle of continuation bytes, which makes us unstoppable in conquering the entire integers, in the name of including them into Unicode codepoints, encoded in UTF-8000.\n              \n\nThe central fascis which sticks out represents the first byte of a code unit. This byte looks different from its continuation bytes, and has the power to start a code unit, which is represented by the wielding of the axe, sticking out on the left. On the right end the central fascis sticks out perhaps representing the terminating 0 of the start bit sequence.\n\nThe favicon for this website is just the Roman numerals VIII.\n\nThe Git repo containing these images is available on GitHub as UTF-8000/UTF-8000-Images.\n\n## Licensing / Copyright (Copyleft)\n\nAs creator of UTF-8000 I want the liberty with which UTF-8000 (the algorithm) can be used to be no less permissible than UTF-8 and ASCII before it. This belongs to everyone. Live Free or Die.\n\nThis website is licensed under CC-BY-NC-SA-4.0, available on GitHub as UTF-8000/UTF-8000-Website.\n\nThe logo / images are licensed under CC-BY-NC-SA-4.0, available on GitHub as UTF-8000/UTF-8000-Images.\n\nThe Python reference implementation is licensed under GPL-3.0-only, available on GitHub as UTF-8000/UTF-8000-Python. Any implementation of decoding and encoding UTF-8000 is going to look somewhat similar to this codebase of course; don&#39;t worry if you want to use MIT or BSD or something else in a clean-room rewrite.\n\n## Thanks\n\nBell Labs:\n\n- \n                  Claude Shannon, founding father of the Digital Age, Information Theory, and Artificial Intelligence. His 1948 paper *A Mathematical Theory of Communication* is the most important mathematics paper of the mid 20th Century, which forms a good chunk of the University of Cambridge Mathematics Tripos course*Coding and Cryptography* . I made a Manim YouTube video in 2024 covering Entropy from Information Theory and its various appearances in mathematics.\n- Ken Thompson, creator of UTF-8, also best known for Unix and other big projects like chess computers.\n- Rob Pike, co-creator of UTF-8, also best known for Plan 9 and other big projects.\n\nThe University of Cambridge mathematics department:\n\n- Professor Stuart Martin, who lectured *Coding and Cryptography* 2019-2020.\n- \n                  Dr Ross Lawther, director of studies for mathematics at Girton College who supervised me for *Coding and Cryptography* (and other mathematics topics!) 2017-2020. It is question 2 of example sheet 1 that concerns the product of two prefix-free codes, which I was reminded of by the product of the self-synchronization and self-punctuation mechanisms of UTF-8000.\n- Dr Keith Carne, whose timeless lecture notes for *Codes and Cryptography* I use often. They are mirrored on my website here. Start at chapter 3 if you&#39;re interested!\n\n## Epilogue\n\n                To think of these stars that you see overhead at night, these vast worlds which we can never reach. I would annex the planets if I could.\n\n- Cecil John Rhodes, founder of Rhodesia.\n\n\nAnd annex the entire integers we have done! I find it fitting that ASCII (🇺🇸) and UTF-8 (🇺🇸) are completed by UTF-8000 (🇬🇧), another Anglosphere classic. But I do have another idea:\n\nI am publishing this document formally on July 4th 2026, perhaps as a 250th birthday gift from Great Britain to the United States of America. Cheers!\n\n                One man alone in a room with a computer, a typewriter as it was, can change the world.\n\n- Jonathan Bowden, English cultural orator.\n\n(As a proud alumnus of Girton College I do feel compelled to comment that man should be read in the Lockean sense of mankind!)\n\n\nThe world has been changed at least twice by Ken Thompson, with Unix and UTF-8. Unix was first written in solitude, in three weeks of the summer of 1969 (the same time as the Moon landing), at Bell Labs on a teletypewriter attached to a spare PDP computer. UTF-8 was invented in one night of autumn 1992, at a New Jersey diner on a placemat, with a slight tweak a few days later.\n\nI, Jay Berry, have so far written this document alone in the summer of 2026, realizing that my reference implementation from autumn 2024 was a nontrivial discovery.</code></pre>","headings":[{"level":1,"text":"UTF-8000","id":"utf-8000"},{"level":2,"text":"TLDR / Examples","id":"tldr-examples"},{"level":2,"text":"Anatomy","id":"anatomy"},{"level":2,"text":"Glossary","id":"glossary"},{"level":2,"text":"Properties","id":"properties"},{"level":3,"text":"Bit Counts","id":"bit-counts"},{"level":3,"text":"Why the Special Cases?","id":"why-the-special-cases"},{"level":3,"text":"Information Rate","id":"information-rate"},{"level":3,"text":"Self-Synchronization","id":"self-synchronization"},{"level":3,"text":"Self-Punctuation","id":"self-punctuation"},{"level":3,"text":"Byte Map","id":"byte-map"},{"level":3,"text":"strcmp(3) Ordering","id":"strcmp-3-ordering"},{"level":3,"text":"No Endianness","id":"no-endianness"},{"level":3,"text":"BOM Support","id":"bom-support"},{"level":3,"text":"Arbitrary Lengths, Sensible Limits","id":"arbitrary-lengths-sensible-limits"},{"level":2,"text":"Intuitive Derivation","id":"intuitive-derivation"},{"level":2,"text":"Encoding","id":"encoding"},{"level":2,"text":"Decoding","id":"decoding"},{"level":3,"text":"Error Recovery","id":"error-recovery"},{"level":3,"text":"The Main Decode Loop","id":"the-main-decode-loop"},{"level":2,"text":"Further Ideas","id":"further-ideas"},{"level":2,"text":"Signed Variant: ZigZag Encoding","id":"signed-variant-zigzag-encoding"},{"level":3,"text":"Source","id":"source"},{"level":3,"text":"TLDR / Examples","id":"tldr-examples-2"},{"level":3,"text":"The ZigZag Function","id":"the-zigzag-function"},{"level":3,"text":"Properties","id":"properties-2"},{"level":3,"text":"Verdict","id":"verdict"},{"level":2,"text":"UTF-16K","id":"utf-16k"},{"level":3,"text":"Source","id":"source-2"},{"level":3,"text":"TLDR / Examples","id":"tldr-examples-3"},{"level":3,"text":"Properties","id":"properties-3"},{"level":3,"text":"Verdict","id":"verdict-2"},{"level":3,"text":"Reference Implementation","id":"reference-implementation"},{"level":2,"text":"UTF-32K","id":"utf-32k"},{"level":3,"text":"Source","id":"source-3"},{"level":3,"text":"TLDR / Examples","id":"tldr-examples-4"},{"level":3,"text":"Properties","id":"properties-4"},{"level":3,"text":"Verdict","id":"verdict-3"},{"level":2,"text":"Rejected Alternatives","id":"rejected-alternatives"},{"level":2,"text":"ASCVI","id":"ascvi"},{"level":3,"text":"Source","id":"source-4"},{"level":3,"text":"TLDR / Examples","id":"tldr-examples-5"},{"level":3,"text":"Properties","id":"properties-5"},{"level":3,"text":"Intuitive Derivation","id":"intuitive-derivation-2"},{"level":3,"text":"Naming","id":"naming"},{"level":3,"text":"Verdict","id":"verdict-4"},{"level":2,"text":"Signed Variant: Two's Complement","id":"signed-variant-two-s-complement"},{"level":3,"text":"Source","id":"source-5"},{"level":3,"text":"TLDR / Examples","id":"tldr-examples-6"},{"level":3,"text":"Properties","id":"properties-6"},{"level":3,"text":"Verdict","id":"verdict-5"},{"level":2,"text":"UTF-Infinity","id":"utf-infinity"},{"level":3,"text":"Source","id":"source-6"},{"level":3,"text":"TLDR / Examples","id":"tldr-examples-7"},{"level":3,"text":"Properties","id":"properties-7"},{"level":3,"text":"Verdict","id":"verdict-6"},{"level":2,"text":"Perl utf8","id":"perl-utf8"},{"level":3,"text":"Source","id":"source-7"},{"level":3,"text":"TLDR / Examples","id":"tldr-examples-8"},{"level":3,"text":"Properties","id":"properties-8"},{"level":3,"text":"Verdict","id":"verdict-7"},{"level":2,"text":"UCS-X","id":"ucs-x"},{"level":3,"text":"Source","id":"source-8"},{"level":3,"text":"TLDR / Examples","id":"tldr-examples-9"},{"level":3,"text":"Properties","id":"properties-9"},{"level":3,"text":"Verdict","id":"verdict-8"},{"level":2,"text":"Owl's Corrected","id":"owl-s-corrected"},{"level":3,"text":"Source","id":"source-9"},{"level":3,"text":"TLDR / Examples","id":"tldr-examples-10"},{"level":3,"text":"Properties","id":"properties-10"},{"level":3,"text":"Verdict","id":"verdict-9"},{"level":2,"text":"Do Nothing","id":"do-nothing"},{"level":3,"text":"Verdict","id":"verdict-10"},{"level":2,"text":"Feedback","id":"feedback"},{"level":3,"text":"Ken Thompson","id":"ken-thompson"},{"level":2,"text":"emails","id":"emails"},{"level":3,"text":"Rejected Alternatives Authors","id":"rejected-alternatives-authors"},{"level":2,"text":"first email","id":"first-email"},{"level":2,"text":"emails","id":"emails-2"}]}}