Language Log — Distribution of Acronym Lengths

Summary: Mark Liberman uses Acronym Finder data to sketch the distribution of initialism lengths, concluding that three-letter initialisms are the most numerous, while four-letter ones decline sharply and five-letter ones are effectively absent.

Sources: Raw/Language Log » Distribution of acronym lengths.md

Last updated: 2026-05-07


Core finding

Liberman sampled random letter strings of lengths 1–5 against Acronym Finder, then extrapolated entry counts:

LengthEstimated entries
1-letter~170
2-letter~3,900
3-letter~83,800
4-letter~64,000
5-letter~0

The peak at three letters is striking. Four-letter strings drop sharply (mean hits per probe falls from ~48 to ~1.4). Five-letter random strings return almost nothing — with occasional exceptions (“backronyms” like DREAM and PATRIOT, which are reverse-engineered from an existing word).

Why initialisms cluster at three letters

Liberman doesn’t theorize directly, but the data suggests a sweet spot: three letters are short enough to be memorable and pronounceable as a unit, long enough to avoid collision with too many other referents. Two-letter strings are scarce because the combinatorial space is small (676 pairs). Four-letter strings are individually rarer because the space is huge (456,976 combinations) and most remain unclaimed.

Complications

  • Acronym vs. initialism: Wiktionary distinguishes initialisms (pronounced letter-by-letter: FBI) from acronyms (pronounced as a word: NATO, laser). Liberman treats them together; the distribution likely differs between the two subtypes.
  • Polysemy is the rule, not the exception: LSA has 123 interpretations in Acronym Finder. In a large corpus (NOW), the Linguistic Society of America gets 55 of 3,680 LSA hits — less than 2%. Context does the disambiguation work that the form cannot.
  • The longest known initialism: MMIWG2SLGBTQQIA+ — 16 characters, from a 2021 Canadian Indigenous rights policy document. A political and cultural artifact, not a linguistic optimum.

Connections

The polysemy of short initialisms is an extreme case of what structure-of-the-lexicon documents generally: the mental lexicon handles massive ambiguity through spreading activation and contextual constraint, not through one-to-one form-meaning mapping.

The corpus methodology here — sampling from a large web corpus (NOW), checking against a reference database — is the standard approach in liberman-pullum-far-from-madding-gerund’s empirical grammatical work.