What counts as one character
JavaScript will tell you a three-person family emoji is eight characters long, and your 280-character limit believes it.
"👨👩👧".length is 8.
That string is one thing on screen, one thing to the cursor, and one thing to the person who typed it. JavaScript reports eight because length counts UTF-16 code units, which is a storage detail that leaked into the language in 1995 and never left.
Three different numbers are defensible depending on what you're asking, and most code picks whichever one is shortest to type.
Three ways to count
Only the last row matches what a person would call one character.
The first row splits surrogate pairs down the middle, so half the entries are unpaired halves that render as nothing. The second row fixes that but still treats the zero-width joiners and skin tone modifiers as separate items. The third row is Intl.Segmenter, and it's the only one that matches what the arrow key does.
Grapheme clusters
A grapheme cluster is the unit a reader thinks of as one character. Sometimes that's one code point. Often it isn't.
Emoji are the obvious case, but they're not the interesting one. Devanagari combines consonants and vowel signs into single conjunct forms. Thai stacks tone marks above and below a base. Hangul composes from jamo. Any Latin text with combining diacritics does it too, which is why "é" can be one code point or two depending on which normalization form it arrived in.
const segmenter = new Intl.Segmenter(undefined, { granularity: 'grapheme' })
const graphemes = [...segmenter.segment(value)].map((entry) => entry.segment)
graphemes.length // what a person would countIntl.Segmenter also does granularity: 'word' and granularity: 'sentence', which is how you split words in Japanese and Thai, neither of which puts spaces between them. Splitting on /\s+/ returns one giant token for an entire Japanese sentence.
Truncation
Counting wrong is a display bug. Cutting wrong is a data bug.
Truncating to a length
Great work 👨…
Great work 👨👩👧 …
That cut landed inside a surrogate pair, so the last character is now an unpaired half.
Slicing at a UTF-16 index can land between the two halves of a surrogate pair. What you store is then a lone surrogate, which is not valid UTF-8, and depending on where it goes next you get a replacement character, a rejected database write, or a JSON parse error two services downstream.
The same applies to the reverse operation. A 280-character limit enforced with .length lets someone post 280 code units, which might be 140 emoji, and rejects a 200-grapheme Hindi sentence that would have fit fine.
Plural rules
The other place English assumptions get compiled in is pluralization.
Plural categories
count === 1 ? singular : plural picks "other" here. This locale needs "few".
Arabic has six plural categories. Polish and Russian have four. Japanese has one, which means 1 file and 2 files are the same word and your English-shaped code inserts a plural that doesn't exist.
count === 1 ? 'item' : 'items' is a hardcoded implementation of English's two-category system. It gets Polish wrong at 2, at 22, and at 102, and gets Arabic wrong at almost every value including zero.
const rules = new Intl.PluralRules(locale)
const label = {
zero: 'no files',
one: '# file',
two: '# files',
few: '# files',
many: '# files',
other: '# files',
}[rules.select(count)]The categories are grammatical names, not numbers. few in Polish means a count ending in 2, 3, or 4 but not ending in 12, 13, or 14. You don't implement that rule. You ask Intl.PluralRules which bucket a number lands in and supply a string per bucket.
The rest of Intl
Intl is a full CLDR database shipped in every browser, and most of it goes uncalled.
Intl.RelativeTimeFormatproduces "3 days ago" in any locale, with correct plural handling, instead of the string concatenation you wroteIntl.ListFormatproduces "a, b, and c" and knows which locales use a serial commaIntl.DisplayNamesturns a language or region code into its name in the user's languageIntl.NumberFormathandles grouping separators, which are periods in German and spaces in FrenchIntl.Collatorsorts strings correctly, because.sort()compares UTF-16 code units and putsZbeforea
None of these need a dependency and all of them are faster than the userland equivalent, because the data is already in the browser.
The character counter is the one worth fixing first. It's the only one on this list where being wrong loses the user's text.