What do you get when you put every character that exists on a single page?
In short: a gigantic HTML file that nearly brings your browser to its knees. And a million chunks of tofu.
Vee Satayamas from Bangkok, Thailand, CC BY 2.0, via Wikimedia Commons
All characters?
First of all, what does all characters mean?
Characters are represented by bytes. And which bytes represent which character is determined by Unicode.
There are over a million Unicode codepoints: 1.114.111 to be exact. In hexadecimal, this is the range U+0000 to U+10FFFF.
According to Wikipedia, Unicode covers 172 scripts. It also includes hyper-specialized things like “Alchemical Symbols”, “Playing Cards”, “Byzantine Musical Symbols” and “Transport and Map Symbols”.
Also, a lot of characters are unasigned. For example, between “Arabic Mathematical Alphabetic Symbols”, ending at U+1EEFF, and the block that follows it, “Mahjong Tiles” starting at U+1F000, there is a gap of 257 characters that don’t have anything assigned to them.
But in the “Arabic Mathematical Alphabetic Symbols” block, which covers 256 characters, only 143 are in use. The other 113 are unused, but reserved for more “Arabic Mathematical Alphabetic Symbols” that might be added in the future.
Anyway, let’s loop over all 1.114.111 characters, put them in an HTML page, and see what happens.
All characters!
If we enclose every unicode value in a <p> tag (the shortest HTML tag without any distracting styling), this results in a 28 megabyte HTML file. Hefty!
It takes Chrome about a minute or two to fully load the page (Eleventy, on the other hand, needs 1.13 seconds to generate it. Go possum!)
Looking at the top of the page, I see mostly familiar characters: punctuation, numbers, the alphabet I know from growing up in Europe. But it quickly shows characters unfamiliar to me, and later on a whole lot of tofu.
Chrome doesn’t like me scrolling through this page and regularly halts, and takes a few seconds when I stop scrolling to recover and decide which characters are currently visible and need to be drawn.
Sprechen sie English?
When the page has finished loading, Chrome decides it’s written in “Chinese (Tradional)” and offers to translate it into English. Presumably this is because Chinese characters are the most abundant in Unicode — no other script has that many individual characters assigned to it.
Letting Chrome translate the page results in some seemingly random characters being replaced by others, or into short text strings like “Get rid of” or “leave”. As expected, the result is useless.
Searching for characters
If I search for a character, the browser highlights all characters that it associates with it. It’s a fuzzy search. Searching for the letter “a” gives 106 hits:
U+0041 A LATIN CAPITAL LETTER A U+0061 a LATIN SMALL LETTER A U+00AA ª FEMININE ORDINAL INDICATOR U+00C0 À LATIN CAPITAL LETTER A WITH GRAVE U+00C1 Á LATIN CAPITAL LETTER A WITH ACUTE U+00C2 Â LATIN CAPITAL LETTER A WITH CIRCUMFLEX U+00C3 Ã LATIN CAPITAL LETTER A WITH TILDE U+00C4 Ä LATIN CAPITAL LETTER A WITH DIAERESIS U+00C5 Å LATIN CAPITAL LETTER A WITH RING ABOVE U+00E0 à LATIN SMALL LETTER A WITH GRAVE U+00E1 á LATIN SMALL LETTER A WITH ACUTE U+00E2 â LATIN SMALL LETTER A WITH CIRCUMFLEX U+00E3 ã LATIN SMALL LETTER A WITH TILDE U+00E4 ä LATIN SMALL LETTER A WITH DIAERESIS U+00E5 å LATIN SMALL LETTER A WITH RING ABOVE U+0100 Ā LATIN CAPITAL LETTER A WITH MACRON U+0101 ā LATIN SMALL LETTER A WITH MACRON U+0102 Ă LATIN CAPITAL LETTER A WITH BREVE U+0103 ă LATIN SMALL LETTER A WITH BREVE U+0104 Ą LATIN CAPITAL LETTER A WITH OGONEK U+0105 ą LATIN SMALL LETTER A WITH OGONEK U+01CD Ǎ LATIN CAPITAL LETTER A WITH CARON U+01CE ǎ LATIN SMALL LETTER A WITH CARON U+01DE Ǟ LATIN CAPITAL LETTER A WITH DIAERESIS AND MACRON U+01DF ǟ LATIN SMALL LETTER A WITH DIAERESIS AND MACRON U+01E0 Ǡ LATIN CAPITAL LETTER A WITH DOT ABOVE AND MACRON U+01E1 ǡ LATIN SMALL LETTER A WITH DOT ABOVE AND MACRON U+01FA Ǻ LATIN CAPITAL LETTER A WITH RING ABOVE AND ACUTE U+01FB ǻ LATIN SMALL LETTER A WITH RING ABOVE AND ACUTE U+0200 Ȁ LATIN CAPITAL LETTER A WITH DOUBLE GRAVE U+0201 ȁ LATIN SMALL LETTER A WITH DOUBLE GRAVE U+0202 Ȃ LATIN CAPITAL LETTER A WITH INVERTED BREVE U+0203 ȃ LATIN SMALL LETTER A WITH INVERTED BREVE U+0226 Ȧ LATIN CAPITAL LETTER A WITH DOT ABOVE U+0227 ȧ LATIN SMALL LETTER A WITH DOT ABOVE U+0363 ͣ COMBINING LATIN SMALL LETTER A U+1D2C ᴬ MODIFIER LETTER CAPITAL A U+1D43 ᵃ MODIFIER LETTER SMALL A U+1DD3 ᷓ COMBINING LATIN SMALL LETTER FLATTENED OPEN A ABOVE U+1DF2 ᷲ COMBINING LATIN SMALL LETTER A WITH DIAERESIS U+1E00 Ḁ LATIN CAPITAL LETTER A WITH RING BELOW U+1E01 ḁ LATIN SMALL LETTER A WITH RING BELOW U+1EA0 Ạ LATIN CAPITAL LETTER A WITH DOT BELOW U+1EA1 ạ LATIN SMALL LETTER A WITH DOT BELOW U+1EA2 Ả LATIN CAPITAL LETTER A WITH HOOK ABOVE U+1EA3 ả LATIN SMALL LETTER A WITH HOOK ABOVE U+1EA4 Ấ LATIN CAPITAL LETTER A WITH CIRCUMFLEX AND ACUTE U+1EA5 ấ LATIN SMALL LETTER A WITH CIRCUMFLEX AND ACUTE U+1EA6 Ầ LATIN CAPITAL LETTER A WITH CIRCUMFLEX AND GRAVE U+1EA7 ầ LATIN SMALL LETTER A WITH CIRCUMFLEX AND GRAVE U+1EA8 Ẩ LATIN CAPITAL LETTER A WITH CIRCUMFLEX AND HOOK ABOVE U+1EA9 ẩ LATIN SMALL LETTER A WITH CIRCUMFLEX AND HOOK ABOVE U+1EAA Ẫ LATIN CAPITAL LETTER A WITH CIRCUMFLEX AND TILDE U+1EAB ẫ LATIN SMALL LETTER A WITH CIRCUMFLEX AND TILDE U+1EAC Ậ LATIN CAPITAL LETTER A WITH CIRCUMFLEX AND DOT BELOW U+1EAD ậ LATIN SMALL LETTER A WITH CIRCUMFLEX AND DOT BELOW U+1EAE Ắ LATIN CAPITAL LETTER A WITH BREVE AND ACUTE U+1EAF ắ LATIN SMALL LETTER A WITH BREVE AND ACUTE U+1EB0 Ằ LATIN CAPITAL LETTER A WITH BREVE AND GRAVE U+1EB1 ằ LATIN SMALL LETTER A WITH BREVE AND GRAVE U+1EB2 Ẳ LATIN CAPITAL LETTER A WITH BREVE AND HOOK ABOVE U+1EB3 ẳ LATIN SMALL LETTER A WITH BREVE AND HOOK ABOVE U+1EB4 Ẵ LATIN CAPITAL LETTER A WITH BREVE AND TILDE U+1EB5 ẵ LATIN SMALL LETTER A WITH BREVE AND TILDE U+1EB6 Ặ LATIN CAPITAL LETTER A WITH BREVE AND DOT BELOW U+1EB7 ặ LATIN SMALL LETTER A WITH BREVE AND DOT BELOW U+2090 ₐ LATIN SUBSCRIPT SMALL LETTER A U+212B Å ANGSTROM SIGN U+24B6 Ⓐ CIRCLED LATIN CAPITAL LETTER A U+24D0 ⓐ CIRCLED LATIN SMALL LETTER A U+A79A Ꞛ LATIN CAPITAL LETTER VOLAPUK AE U+A79B ꞛ LATIN SMALL LETTER VOLAPUK AE U+A7C0 Ꟁ LATIN CAPITAL LETTER OLD POLISH O U+A7C1 ꟁ LATIN SMALL LETTER OLD POLISH O U+FF21 A FULLWIDTH LATIN CAPITAL LETTER A U+FF41 a FULLWIDTH LATIN SMALL LETTER A U+1CCD6 OUTLINED LATIN CAPITAL LETTER A U+1D400 𝐀 MATHEMATICAL BOLD CAPITAL A U+1D41A 𝐚 MATHEMATICAL BOLD SMALL A U+1D434 𝐴 MATHEMATICAL ITALIC CAPITAL A U+1D44E 𝑎 MATHEMATICAL ITALIC SMALL A U+1D468 𝑨 MATHEMATICAL BOLD ITALIC CAPITAL A U+1D482 𝒂 MATHEMATICAL BOLD ITALIC SMALL A U+1D49C 𝒜 MATHEMATICAL SCRIPT CAPITAL A U+1D4B6 𝒶 MATHEMATICAL SCRIPT SMALL A U+1D4D0 𝓐 MATHEMATICAL BOLD SCRIPT CAPITAL A U+1D4EA 𝓪 MATHEMATICAL BOLD SCRIPT SMALL A U+1D504 𝔄 MATHEMATICAL FRAKTUR CAPITAL A U+1D51E 𝔞 MATHEMATICAL FRAKTUR SMALL A U+1D538 𝔸 MATHEMATICAL DOUBLE-STRUCK CAPITAL A U+1D552 𝕒 MATHEMATICAL DOUBLE-STRUCK SMALL A U+1D56C 𝕬 MATHEMATICAL BOLD FRAKTUR CAPITAL A U+1D586 𝖆 MATHEMATICAL BOLD FRAKTUR SMALL A U+1D5A0 𝖠 MATHEMATICAL SANS-SERIF CAPITAL A U+1D5BA 𝖺 MATHEMATICAL SANS-SERIF SMALL A U+1D5D4 𝗔 MATHEMATICAL SANS-SERIF BOLD CAPITAL A U+1D5EE 𝗮 MATHEMATICAL SANS-SERIF BOLD SMALL A U+1D608 𝘈 MATHEMATICAL SANS-SERIF ITALIC CAPITAL A U+1D622 𝘢 MATHEMATICAL SANS-SERIF ITALIC SMALL A U+1D63C 𝘼 MATHEMATICAL SANS-SERIF BOLD ITALIC CAPITAL A U+1D656 𝙖 MATHEMATICAL SANS-SERIF BOLD ITALIC SMALL A U+1D670 𝙰 MATHEMATICAL MONOSPACE CAPITAL A U+1D68A 𝚊 MATHEMATICAL MONOSPACE SMALL A U+1F130 🄰 SQUARED LATIN CAPITAL LETTER A U+1F150 🅐 NEGATIVE CIRCLED LATIN CAPITAL LETTER A U+1F170 🅰 NEGATIVE SQUARED LATIN CAPITAL LETTER A
It finds “a” regardless of casing or diacritics, which is not that surpising. But it also find all kinds of other characters that could be considerd an “a”, like ⓐ or 𝖆. Neat!
Which fonts do you need to render every character that exists?
Chrome on my Macbook uses 167 different fonts to render this page.
The page is rendered without any font-family declared in CSS, so the browser defaults to Times New Roman (or does it?! — read on!). When Times New Roman doesn’t contain a certain character, it asks the OS for a font that does. For example, “Songti SC” is the next most used font on the list, which is used for (most of) the Chinese characters on the page.
This is the full list of fonts used:
1041000 Times 31542 Songti SC 11723 AppleMyungjo 4565 .蘋方-港 2190 Lucida Grande 1927 Apple Symbols 1314 Apple Color Emoji 1234 Noto Sans Cuneiform 1232 Hiragino Mincho ProN 1161 STIX Two Math 1071 Noto Sans EgyptHiero 1033 Geeza Pro 1003 Times New Roman 726 Euphemia UCAS 657 Noto Sans Bamum 496 Kefa III 381 Tamil Sangam MN 341 Noto Sans Linear A 324 Menlo 300 Noto Sans Vai 287 Noto Serif Myanmar 278 Arial Unicode MS 263 .Hiragino Sans GB Interface W3 256 Apple Braille 220 Kokonor 213 Noto Sans Mende Kikakui 211 Noto Sans Linear B 209 Noto Sans Miao 179 .蘋方-繁 176 Galvji 172 Geneva 168 Zapf Dingbats 165 Noto Serif Balinese 160 Helvetica 158 ITF Devanagari 151 Noto Sans Coptic 149 Noto Sans Duployan 145 Khmer MN 134 Bradley Hand 134 Noto Sans Pahawh Hmong 132 Noto Sans Glagolitic 127 Noto Sans Tai Tham 122 Noto Sans Sharada 122 Noto Sans Meroitic 121 Kohinoor Telugu 118 Kohinoor Gujarati 115 Kohinoor Bangla 114 Noto Sans Newa 114 Noto Sans Siddham 113 Noto Sans Bhaiksuki 109 MuktaMahee Regular 109 Noto Sans Javanese 108 Noto Sans OldHung 106 Noto Sans Tirhuta 104 Noto Sans Marchen 102 Noto Sans Saurashtra 100 Malayalam MN 100 Noto Sans Cham 96 Noto Sans Adlam 96 Noto Sans MeeteiMayek 96 Noto Sans Modi 94 Noto Sans Armenian 94 Kannada MN 94 Noto Sans Lepcha 93 Noto Sans Masaram Gondi 92 Noto Sans Limbu 91 Noto Sans Chakma 90 Oriya MN 88 Noto Sans Sundanese 87 Thonburi 84 Noto Sans WarangCiti 83 Noto Sans NewTaiLue 83 Noto Sans Kaithi 82 Noto Sans Syriac 81 Noto Sans Tai Viet 81 Noto Sans Khudawadi 80 Sinhala MN 80 Noto Sans Takri 78 Noto Sans Kharoshthi 75 Noto Sans Khojki 75 Noto Sans Gunjala Gondi 73 Noto Sans Old Turkic 73 Noto Serif Ahom 72 Noto Sans NKo 72 Noto Sans Osage 71 Noto Serif Hmong Nyiakeng 70 Noto Sans Batak 67 .SF Arabic 65 Lao MN 65 Noto Sans CaucAlban 63 Noto Sans Wancho 61 Noto Sans Samaritan 61 Noto Sans Avestan 60 Noto Sans Tifinagh 58 Songti TC 57 Noto Sans PauCinHau 56 Noto Sans PhagsPa 55 Noto Sans Kayah Li 55 Noto Sans Cypriot 54 Noto Sans HanifiRohg 53 Noto Sans Manichaean 50 Noto Sans Thaana 50 Noto Sans Rejang 50 Noto Sans OldPersian 49 Noto Sans Carian 48 Noto Sans Ol Chiki 48 .CJK Symbols Fallback HK 48 Noto Sans Lisu 47 Noto Serif Yezidi 46 Noto Sans Nag Mundari 45 Noto Sans Syloti Nagri 43 Noto Sans Old Permic 43 Noto Sans Mro 40 Noto Sans Osmanya 40 Noto Sans Elbasan 40 Noto Sans Nabataean 40 Noto Sans Mahajani 39 Noto Sans Old Italic 38 Noto Sans Multani 36 Noto Sans Bassa Vah 35 Noto Sans Tai Le 35 Noto Sans Buginese 35 Noto Sans SoraSomp 32 Noto Sans Mandaic 32 Noto Sans Palmyrene 32 Noto Sans OldSouArab 32 Noto Sans OldNorArab 31 Sickness 31 Noto Sans Ugaritic 31 Noto Sans ImpAramaic 30 Noto Sans InsParthi 29 Noto Sans Lycian 29 Noto Sans Phoenician 29 Noto Sans PsaPahlavi 27 Noto Sans Tagalog 27 Noto Sans Gothic 27 Noto Sans Lydian 27 Noto Sans InsPahlavi 26 Noto Sans Hanunoo 26 Noto Sans Hatran 22 Noto Sans Buhid 22 .蘋方-簡 21 .SF Malayalam 20 Noto Nastaliq Urdu 20 Noto Sans Tagbanwa 17 Big Caslon 14 DIN Condensed 13 .SF NS 13 Noto Sans Mongolian 12 .New York 12 .SF Gujarati 8 Baskerville 7 Athelas 6 .SF Telugu 4 .Apple SD Gothic NeoI 3 .SF Devanagari 2 .SF Odia 2 .SF Kannada 2 Noto Sans Kannada 2 Rockwell 2 .SF Compact 2 Helvetica Neue 1 Charter 1 .DecoType Nastaleeq Urdu UI 1 Sukhumvit Set 1 Noto Sans Yi 1 Ayuthaya
Times (New Roman)
Note that “Times” and “Times New Roman” are both on the list.
Turns out that Times is the main font, the alpha supreme first lady of my computer’s font stack. The “A” is rendered in Times, and so is the rest of the alphabet.
But there’s also Times New Roman. Isn’t that the same? Isn’t that the one we always use for unstyled text? Well, apparently not. Times is used for the more common Latin characters at the very start of Unicode, up to U+017F. Times New Roman is used for 1003 characters starting at “Latin Extended-B” (U+0180) and beyond.
Another observation: even though Times New Roman contains Arabic characters (appearing after U+0180) and thus technically supports Arabic, it is not being used for Arabic. Instead, the browser uses Geeza Pro. Presumably this is because Geeza Pro is a better font. Better looking, more characters, more weights — who knows, but my computer prefers it to render Arabic over Times New Roman.
But, back to Times. Times doesn’t support 1.041.000 characters — it supports 1.284. So why is this number so large?
Tofu or not tofu
Tofu is the square box you see when a character can’t be rendered by a font. It’s a fallback that is shown to communicate “we can’t show this character!”
And this page has A. LOT. OF. TOFU.
Turns out that when a character can’t be rendered, the browser asks the first font in the stack for a tofu character to show instead. So all tofus are rendered using Times!
On my system, the first Unicode value that can’t be rendered, and thus shows a tofu, is U+0378.
According to Wikipedia, U+0378 should be part of the “Greek and Coptic” Unicode block, which goes from U+0370 to U+03FF, which is 144 codepoints. But only 135 of those have a character assigned, leaving 9 characters without a valid aignment.
The Wikipedia article talks about characters being withdrawn or replaced, so I’m assuming it got added, someone went “You know what, nah” and removed it. Or maybe it wasn’t assigned to anything to begin with, for reasons unknown (to me).
It looks like most of the tofu comes from these highly specialized Unicode blocks, so exotic and strange that no OS bothers to ship fonts for it — if these fonts even exist at all.
In my own area of interest, for example, “Symbols for Legacy Computing” sounds fun and includes a lot of neat and bizarre characters. But my OS doesn’t have a font for it, and the number of fonts out there supporting this Unicode range is very small. So they’re all tofu.
Counting tofu
So we have a lot of tofu. But exactly how much is “a lot”?
If all 1.284 characters of Times are used on the page, that means the 1.039.716 remaining Times characters use Times just for its tofu.
But counting all characters rendered in the other fonts, we get to 74.012 proper characters. With the 1.284 proper Times characters added (assuming they’re all used, which is probably a wrong assumption anyway), that’d mean that 75.296 characters are not tofu, meaning there should be 1.038.815 tofu.
A third check, rendering everything to canvas and doing a visual search for the tofu character, results in 78.576 proper characters and 1.035.535 tofu.
In other words, just like it’s impossible to know the number of tofus in the real world, it’s impossible to know how many there are on the test page. But, like in the real world, it’s safe to say there are more than a million — roughly 90% of all Unicode characters.
Least used fonts
Back to our list of 167 fonts. Which characters are rendered by the fonts at the bottom of the list?
Baskerville `U+1EFA` Ỻ LATIN CAPITAL LETTER MIDDLE-WELSH LL `U+1EFB` ỻ LATIN SMALL LETTER MIDDLE-WELSH LL `U+1EFC` Ỽ LATIN CAPITAL LETTER MIDDLE-WELSH V `U+1EFD` ỽ LATIN SMALL LETTER MIDDLE-WELSH V `U+1EFE` Ỿ LATIN CAPITAL LETTER Y WITH LOOP `U+1EFF` ỿ LATIN SMALL LETTER Y WITH LOOP `U+2C6D` Ɑ LATIN CAPITAL LETTER ALPHA `U+2C70` Ɒ LATIN CAPITAL LETTER TURNED ALPHA Athelas `U+0370` Ͱ GREEK CAPITAL LETTER HETA `U+0371` ͱ GREEK SMALL LETTER HETA `U+0372` Ͳ GREEK CAPITAL LETTER ARCHAIC SAMPI `U+0373` ͳ GREEK SMALL LETTER ARCHAIC SAMPI `U+0376` Ͷ GREEK CAPITAL LETTER PAMPHYLIAN DIGAMMA `U+0377` ͷ GREEK SMALL LETTER PAMPHYLIAN DIGAMMA `U+03CF` Ϗ GREEK CAPITAL KAI SYMBOL .SF Telugu `U+0C04` ఄ TELUGU SIGN COMBINING ANUSVARA ABOVE `U+0C3C` ఼ TELUGU SIGN NUKTA `U+0C5D` ౝ TELUGU LETTER NAKAARA POLLU `U+0C77` ౷ TELUGU SIGN SIDDHAM .Apple SD Gothic NeoI `U+115F` ᅟ HANGUL CHOSEONG FILLER `U+1160` ᅠ HANGUL JUNGSEONG FILLER `U+119E` ᆞ HANGUL JUNGSEONG ARAEA `U+11A2` ᆢ HANGUL JUNGSEONG SSANGARAEA .SF Devanagari `U+A8FE` ꣾ DEVANAGARI LETTER AY `U+A8FF` ꣿ DEVANAGARI VOWEL SIGN AY .SF Odia `U+0B55` ୕ ORIYA SIGN OVERLINE .SF Kannada `U+0C84` ಄ KANNADA SIGN SIDDHAM `U+0CDD` ೝ KANNADA LETTER NAKAARA POLLU Noto Sans Kannada `U+1CDA` ᳚ VEDIC TONE DOUBLE SVARITA `U+1CF5` ᳵ VEDIC SIGN JIHVAMULIYA Rockwell `U+20B7` ₷ SPESMILO SIGN `U+20BB` ₻ NORDIC MARK SIGN .SF Compact `U+20C1` UNKNOWN `U+2B58` ⭘ HEAVY CIRCLE Helvetica Neue `U+2B90` ⮐ RETURN LEFT `U+2B91` ⮑ RETURN RIGHT Charter `U+0080` Control character .DecoType Nastaleeq Urdu UI `U+061D` ؝ ARABIC END OF TEXT MARK Sukhumvit Set `U+0E00` UNKNOWN Noto Sans Yi `U+A2A1` ꊡ YI SYLLABLE ZEP Ayuthaya `U+FB0F` UNKNOWN
Ayuthaya, at the bottom of our list, supports Latin, Cyrillic and Thai. But it’s only used for one single character: U+FB0F. This is an “Undefined Character” in the “Basic Multilingual Plane”.
Likewise, Sukhumvit Set is used for U+0E00, another undefined character.
It seems these fonts are somehow chosen to render these uncommon or unused codepoints. The left-overs of the Unicode range. Why? Didn’t any of the roughly 160 fonts earlier in the stack do a good enough job?
Are these really all the characters that exist?
No!
There exist more characters than are currently in Unicode. A lot of obscure scripts and languages are missing. Or they are yet incomplete, and being expanded.
There are also constantly characters being suggested and rejected. I couldn’t find a list of recent rejections, other than this interesting list of rejected emoji proposals.
But even within the currently “existing” characters, more characters can be made by combining Unicode characters. There are practical uses for that, but also emoji use the zero width joiner to combine emoji with different meanings into a new emoji. For example, a guy combined with a rocket is an astronaut: 👨 + 🚀 = 👨🚀.
So, even with 1.114.111 characters to your disposal, you can put even more characters on the screen!
Okay, so what did we get when we put every character that exists on a single page?
For me, fun. I got to slap together the sort of code where you get to name variables with whatever pops in your head. So yes, fnurk holds the number of tofu found by window.find — whatcha gonna do about it? It was stupid, carefree hacking, yielding useless results even by the most optimistic expectations.
And, just as satisfying, I also did learn a thing or two:
- Most characters you read on an unstyled HTML page aren’t in Times New Roman, they’re in Times.
- There’s a Unicode block for an ancient, undeciphered script.
- text.makeup knows about every single character out there.
- Somebody had to make a font stack of 167 fonts to take care of all the fallbacks. Think of that when you add
sans-serifto yourfont-familyand call it a day! - Unicode, languages, scripts and all the squiggly drawings that humans created to communicate ideas with each other are immensly complex and contain a wealth of human history. Some of these rabbit holes go so deep they come out the other end, then continue into space.
- There is a codepoint for Pac-Man!