What counts as one character

JavaScript will tell you a three-person family emoji is eight characters long, and your 280-character limit believes it.

"👨‍👩‍👧".length is 8.

That string is one thing on screen, one thing to the cursor, and one thing to the person who typed it. JavaScript reports eight because length counts UTF-16 code units, which is a storage detail that leaked into the language in 1995 and never left.

Three different numbers are defensible depending on what you're asking, and most code picks whichever one is shortest to type.

Three ways to count

string.length8
�U+D83D�U+DC68‍U+200D�U+D83D�U+DC69‍U+200D�U+D83D�U+DC67
UTF-16 code units. Surrogate pairs are split down the middle.
[...string]5
👨U+1F468‍U+200D👩U+1F469‍U+200D👧U+1F467
Code points. Joiners and modifiers still count as separate items.
Intl.Segmenter1
👨‍👩‍👧U+1F468
Grapheme clusters. What the cursor moves over when you press the arrow key.

Only the last row matches what a person would call one character.

The first row splits surrogate pairs down the middle, so half the entries are unpaired halves that render as nothing. The second row fixes that but still treats the zero-width joiners and skin tone modifiers as separate items. The third row is Intl.Segmenter, and it's the only one that matches what the arrow key does.

Grapheme clusters

A grapheme cluster is the unit a reader thinks of as one character. Sometimes that's one code point. Often it isn't.

Emoji are the obvious case, but they're not the interesting one. Devanagari combines consonants and vowel signs into single conjunct forms. Thai stacks tone marks above and below a base. Hangul composes from jamo. Any Latin text with combining diacritics does it too, which is why "é" can be one code point or two depending on which normalization form it arrived in.

const segmenter = new Intl.Segmenter(undefined, { granularity: 'grapheme' })
const graphemes = [...segmenter.segment(value)].map((entry) => entry.segment)
 
graphemes.length // what a person would count

Intl.Segmenter also does granularity: 'word' and granularity: 'sentence', which is how you split words in Japanese and Thai, neither of which puts spaces between them. Splitting on /\s+/ returns one giant token for an entire Japanese sentence.

Truncation

Counting wrong is a display bug. Cutting wrong is a data bug.

Truncating to a length

value.slice(0, 13)

Great work 👨…

first 13 graphemes

Great work 👨‍👩‍👧 …

That cut landed inside a surrogate pair, so the last character is now an unpaired half.

Slicing at a UTF-16 index can land between the two halves of a surrogate pair. What you store is then a lone surrogate, which is not valid UTF-8, and depending on where it goes next you get a replacement character, a rejected database write, or a JSON parse error two services downstream.

The same applies to the reverse operation. A 280-character limit enforced with .length lets someone post 280 code units, which might be 140 emoji, and rejects a 200-grapheme Hindi sentence that would have fit fine.

Plural rules

The other place English assumptions get compiled in is pluralization.

Plural categories

zeroonetwofewmanyother
ar has 6 formsnew Intl.PluralRules('ar').select(3) === 'few'

count === 1 ? singular : plural picks "other" here. This locale needs "few".

Arabic has six plural categories. Polish and Russian have four. Japanese has one, which means 1 file and 2 files are the same word and your English-shaped code inserts a plural that doesn't exist.

count === 1 ? 'item' : 'items' is a hardcoded implementation of English's two-category system. It gets Polish wrong at 2, at 22, and at 102, and gets Arabic wrong at almost every value including zero.

const rules = new Intl.PluralRules(locale)
 
const label = {
  zero: 'no files',
  one: '# file',
  two: '# files',
  few: '# files',
  many: '# files',
  other: '# files',
}[rules.select(count)]

The categories are grammatical names, not numbers. few in Polish means a count ending in 2, 3, or 4 but not ending in 12, 13, or 14. You don't implement that rule. You ask Intl.PluralRules which bucket a number lands in and supply a string per bucket.

The rest of Intl

Intl is a full CLDR database shipped in every browser, and most of it goes uncalled.

  • Intl.RelativeTimeFormat produces "3 days ago" in any locale, with correct plural handling, instead of the string concatenation you wrote
  • Intl.ListFormat produces "a, b, and c" and knows which locales use a serial comma
  • Intl.DisplayNames turns a language or region code into its name in the user's language
  • Intl.NumberFormat handles grouping separators, which are periods in German and spaces in French
  • Intl.Collator sorts strings correctly, because .sort() compares UTF-16 code units and puts Z before a

None of these need a dependency and all of them are faster than the userland equivalent, because the data is already in the browser.

The character counter is the one worth fixing first. It's the only one on this list where being wrong loses the user's text.