Original research · 2026-06-02
English Initial Letter Distribution: Which Letters Generate the Most Words and Confusables (2026)
Not all 26 letters share the vocabulary burden equally. The distributions below reflect the current dataset, two separate questions with notably different answers, each driven by distinct structural forces in English morphology.
Two distributions, two questions
This article presents two parallel letter-frequency analyses over different working cohorts. The first counts frequency-ranked English entries used by PlainSpell's browse surface. The second counts filtered confusable pairs whose two members are frequency-ranked, at most rank 30,000, at least four letters long, and similar in length. These are useful discovery views, not counts of every Wiktionary entry or every possible pair.
Both distributions derive from the PlainSpell English corpus, which is sourced from Wiktionary. Letter counts reflect the first Unicode character of the headword after lowercasing and normalization, within the filters stated above.
Chart 1: Top 12 letters by vocabulary entry count
English words by initial letter
Top 12 initial letters among frequency-ranked English entries used by the browse surface
- S
S
6,566 words
- C
C
5,599 words
- P
P
4,417 words
- B
B
4,121 words
- M
M
4,057 words
- A
A
3,931 words
- D
D
3,470 words
- T
T
3,151 words
- R
R
3,094 words
- H
H
2,646 words
- F 2,507
F
2,507 words
- G 2,313
G
2,313 words
Browse the vocabulary behind each letter: s · c · p · b · m · a · d · t · r · h · f · g
Chart 2: Top 12 letters by confusable pair count
Confusable pairs by initial letter
Top 12 initial letters among filtered, frequency-ranked confusable pairs
- S
S
23,373 confusable pairs
- C
C
16,243 confusable pairs
- B
B
14,230 confusable pairs
- M
M
10,419 confusable pairs
- P
P
9,787 confusable pairs
- T 7,998
T
7,998 confusable pairs
- L 7,624
L
7,624 confusable pairs
- D 7,489
D
7,489 confusable pairs
- R 7,263
R
7,263 confusable pairs
- F 6,770
F
6,770 confusable pairs
- H 6,407
H
6,407 confusable pairs
- A 5,140
A
5,140 confusable pairs
Finding 1: Latin and Greek prefixes cluster under 's', 'p', and 'c'
The high entry counts for the letters s, p, and c in English vocabulary are not accidental, they reflect the productive prefix morphology that English inherited from Latin and Greek through successive waves of scholarly and ecclesiastical borrowing. The Latin prefixes sub-, super-, semi-, syn-, sym- all begin with s; pre-, pro-, per-, post-, para- all begin with p; and con-, com-, contra-, circum-, cata- all begin with c. Each of these prefixes is attached to hundreds of stems, generating large families of derived words that all share an initial letter. The result is a heavily skewed distribution where three letters account for a disproportionate share of the total vocabulary. Old English native roots (the short, high-frequency function words) contribute relatively little to total entry count, they are few in number and deeply familiar, while the vast Latinate and Greek scientific vocabulary dominates the count, and that vocabulary is prefix-organized under precisely these three letters.
Finding 2: 's', 'c', and 'b' lead confusable pairs for different reasons
The confusable pair leadership of s, c, and b does not have a single explanation. For s, the cause is similar to the word-count leadership: a large vocabulary under s means more opportunities for edit-distance proximity between distinct entries. Many s-initial words are prefix variants of each other (sub/sup, syn/sym) where the prefix alternation produces a real confusable pair. For c, the con-/com- prefix family creates dense clusters: conform and comfort, contest and context, conscience and conscious are all near-matches generated by morphological productivity under the same prefix. The letter b tells a different story: it benefits from a high density of short native English words (be, by, but, bid, bad, bag, ban, bar, bat, bay) that form a tightly interconnected edit-distance network. The short-word effect described in the confusable-pairs research article applies with particular force here, a 3-letter b-word has an unusually large fraction of edit-distance-1 neighbors that are themselves real English words.
Finding 3: The 'un-' and 're-' prefix families boost 'u' and 'r' beyond their baseline
Two letters whose vocabulary counts exceed what letter-frequency in ordinary English text would predict are u and r. The productive English prefixes un- (negation of adjectives and verbs: unclear, uneven, unwilling, unreliable) and re- (repetition: rebuild, reconsider, rearrange, redefine) generate large derived-word families under these initial letters. The un- family is particularly large because English applies it to an almost unlimited range of adjectives and past participles, while re- is equally productive with verbs. Neither prefix family has an especially close Latinate rival, dis- and de- are the nearest competitors for negation and reversal, so the u and r counts reflect genuine morphological productivity rather than a borrowed prefix cluster.
What this means for spelling-resource coverage and writer guidance
Dictionary editors and spelling-resource designers who aim for comprehensive coverage face an uneven workload: adding entries under s, p, and c requires processing substantially more candidate words than adding entries under x, z, or q. The entry-count distribution shown above is a rough proxy for editorial effort allocation. It also has practical implications for writers. If you are writing in a domain that makes heavy use of Latin-derived technical vocabulary, medicine, law, science, policy, the letters s, p, and c are where your highest-risk misspellings and confusable pairs will cluster. Proofreading strategies that treat all 26 letters as equally likely sites of error are miscalibrated for this reality. A targeted review of s-initial, p-initial, and c-initial technical terms before submission is a higher-ROI activity than a uniform word-by-word check. The full methodology used to compute these distributions, including how compound words, hyphenated forms, and proper nouns are handled, is documented on the PlainSpell methodology page.
Letter distribution across languages: a brief comparative note
The s/p/c dominance pattern is specific to English and its particular Latin-Greek inheritance. French shows a similar but more pronounced s and p bias, reflecting the same prefix families reinforced by Romance inflectional morphology. German's letter distribution is flatter and more spread across the alphabet, because German productive morphology relies heavily on compounding (which distributes first-letter counts more evenly) rather than prefix derivation (which concentrates them). Spanish, like French, shows strong s and p counts. Portuguese is broadly similar to Spanish in distribution shape but with smaller absolute counts due to the corpus coverage gap documented in the Wiktionary coverage analysis. Cross-language letter distributions are available in the per-language browse sections of PlainSpell.
Methodology
The word distribution extracts the first character (lowercased, Unicode-normalized) from entries with a non-null frequency rank. The confusable distribution is narrower: both members must have a frequency rank at or above the working threshold, be at least four letters long, and differ in length by no more than three characters. Only ASCII a-z entries are counted; entries beginning with digits, punctuation, or non-ASCII characters are excluded. The full specification is on the PlainSpell methodology page.
Limitations: Counting only within-initial-letter confusable pairs means the confusable_letter_counts table undercounts true cross-letter confusable pairs (e.g. "cue" vs "queue" share an initial letter by our rule but differ in standard alphabetic ordering). The within-letter restriction is a computational simplification. Words beginning with capital letters (proper nouns) are lowercased and included in the same letter bucket, which slightly inflates counts for letters with common proper-noun initials (M for Mary/Mark/March, J for John/June/July).
Sources
Source: Wiktionary (English edition) JSONL dump via wiktextract · 2026 Open data under CC BY-SA 4.0.
Source: Bauer, Laurie, English Word-Formation Cambridge University Press · 1983 Foundational reference on English prefix and suffix morphology.
Source: Crystal, David, The Cambridge Encyclopedia of the English Language 2nd Edition · 2003 Reference work on English vocabulary history and morphological structure.