How many Chinese characters do you actually need?
Ask this question anywhere and you get the same answer: two to three thousand. It gets repeated because it is roughly right about one thing and quietly wrong about another. It is close to what near-complete character coverage costs, the point where almost every character on a page is one you have seen before. It is a terrible answer to what most people are actually asking, which is how many they need before Chinese starts being readable at all.
A single number cannot answer that, because character frequency is extraordinarily lopsided. So instead of repeating the folklore, we counted. What follows is a coverage curve built from 33 billion characters of real Chinese, spoken and written, showing what share of an ordinary page the most common 100, 500, 1,000 or 3,000 characters actually account for.
The short version
- ~170 characters covers half of everything you read
- ~510 characters covers 75%
- ~1,090 characters covers 90%
- ~2,310 characters covers 98%
- The first 500 do three quarters of the work. Characters 2,501-3,000 add 0.7%.
Where this comes from: we make Hanzio, a Chinese dictionary and study app, so we keep frequency data for the whole language anyway. The numbers below come from that data - the same data Hanzio's own frequency scores are built from - not from a published estimate. The full method is in How we counted, including the parts that would make these figures wrong for you.
How many characters do you need?
About 1,000 characters covers roughly 90% of the characters in everyday Chinese text. 500 covers about 75%. To reach 98% you need around 2,300, and 99% costs around 2,900. Half of everything you will ever read is carried by just 170 characters.
| Characters known | Share of running text | Unknown characters per 100 |
|---|---|---|
| 100 | 40% | 60 |
| 250 | 59% | 41 |
| 500 | 75% | 25 |
| 1,000 | 89% | 11 |
| 1,500 | 94% | 6 |
| 2,000 | 97% | 3 |
| 2,500 | 98% | 2 |
| 3,000 | 99% | 1 |
Read that last column before the middle one. It is the same information, and it is much harder to feel good about.
What that number does not mean
Coverage percentages are seductive and they are routinely oversold, including by people selling Chinese apps. Two things have to be said before the interesting part.
Coverage is not comprehension
At 1,000 characters you know about 89% of the characters on a page. That is not 89% of the meaning. It is eleven unknown characters in every hundred - roughly one per line, and they will not be evenly spread. They cluster in exactly the places that carry the information: names, technical terms, the specific verb the sentence turns on.
Even 97%, which takes 2,000 characters, is three unknowns per hundred. A page of a novel holds a few hundred characters, so that is a handful of gaps per page. This is workable reading with a dictionary. It is not comfortable reading. Anyone telling you a four-digit character count makes Chinese readable is rounding a real curve into a slogan.
(There is a well-known finding in second-language reading research that unassisted comprehension needs around 98% coverage. It is worth knowing, but it is about word coverage, not characters, so it does not transfer directly to the table above. We mention it as an analogy, not as evidence.)
Knowing characters is not knowing words
Chinese words are mostly built from two or more characters, and the combination is often not predictable from the parts. 意 and 思 are both common and individually unremarkable; 意思 means "meaning". 电 is electricity and 脑 is brain; 电脑 is a computer. You can know every character in a sentence and still not know the sentence.
So characters are a good unit to attack first, not a substitute for vocabulary. They are simply much more efficient per item at the start: reaching that same 89% coverage with whole words instead takes around 11,800 words. (Those two figures are measured in different units - 89% of character tokens and 89% of word tokens - so treat "about twelve times as many" as the shape of the difference rather than a precise ratio.)
The coverage curve
Here is the whole thing. The horizontal axis is how many characters you know, most common first; the vertical axis is the share of the characters on an ordinary page that those cover. Coverage, not comprehension - see above.
Share of running text covered by the most common characters
Two features matter. The curve is almost vertical at the left edge, and it is almost horizontal at the right. Nearly everything useful about learning Chinese characters follows from those two facts.
The first 500 do most of the work
Split the first 3,000 characters into blocks of 500 and ask what each block buys you.
| 1-500 | +74.6 | |
|---|---|---|
| 501-1,000 | +14.1 | |
| 1,001-1,500 | +5.6 | |
| 1,501-2,000 | +2.7 | |
| 2,001-2,500 | +1.4 | |
| 2,501-3,000 | +0.7 |
The first block is worth more than the other five put together, several times over. And it concentrates further down than that: the thirty most common characters on their own account for about 23% of all running text. Thirty. These ones:
的我是了不在你有一人这个们大来他要上中为到会和好就国么以时生
Nothing about that list is surprising. It is mostly grammar, pronouns and the handful of verbs every sentence needs, which is exactly the point. The characters doing the heaviest lifting in Chinese are structural and unavoidable, and you will meet them on day one whether anyone plans it or not.
The last thousand buys almost nothing
Now read the same table from the bottom. Characters 2,001 to 3,000 - a thousand characters, which for most learners is a year or more of steady work - add about two percentage points between them. The thirty characters above are worth more than ten times that.
This is not an argument for stopping at 2,000. Those last characters are what take you from looking something up every line or two to looking something up once a paragraph: the gap between 97% and 99% is the difference between three unknown characters per hundred and one. Fewer interruptions, not none - the vocabulary those characters spell out is a separate job, and it is the one that decides whether you actually understand the page.
It is an argument about order, and about what to expect. The early stretch pays out fast enough that progress feels dramatic. The later stretch does not, and learners routinely read that slowdown as their own failure rather than as the shape of the distribution. It is the shape of the distribution.
It depends a little on what you read
The curve above blends registers, so it describes a general reader rather than any specific one. Splitting them apart shifts the numbers, though less than you might expect, and not always in the obvious direction.
Dialogue is the easiest by some distance: 90% coverage arrives at about 810 characters against 1,090 for the blend, and that lead widens to roughly 540 characters by the time you reach 99%. Speech is not simply a smaller sample, either - it uses more distinct characters than most written material does, and still packs more of its text into the commonest few hundred. It is not that people talking have a smaller vocabulary; it is that they lean much harder on the top of it. Written material varies too, but over a narrower band: roughly 150 to 200 characters between the easiest and the hardest at 90% coverage, with the more formulaic the writing, the earlier the curve arrives.
The size of that effect is a few hundred characters either side of the figures above. It does not move you from 1,000 to 3,000. Read a lot of prose and expect the curve to run slightly behind; lean on dialogue and it runs ahead.
Where "2,000-3,000" comes from
The folklore number is not invented, and it is not really wrong. It answers a different question.
The Table of General Standard Chinese Characters lists 3,500 characters at its first level, described as covering the needs of everyday life, and 8,105 in total. Educated adult native speakers are usually estimated at around 6,000. Those are inventory counts: how many characters a standard contains, or a person has accumulated over a lifetime.
Coverage of running text is a different measure, and on that measure 2,000-3,000 lands in the 97-99% band. So the number quietly describes what near-complete character coverage costs. It gets quoted to beginners as though it were the entry price, and that is the only thing wrong with it - nobody mentions that most of the value arrives in the first fifth of the trip.
What this means for what you study
Three things follow, and none of them is complicated.
Frequency order beats textbook order for the first thousand characters. While the curve is steep, the difference between a well-ordered list and a badly-ordered one is enormous - hundreds of characters of wasted effort. Once you are past 1,500 the curve is flat enough that order stops mattering much, and what you read should drive what you learn instead.
Learn characters inside words, not as a list. The efficiency argument for characters is real, but a character learned in isolation is a shape with a gloss attached. Learned inside 电脑 or 意思, it comes with the thing you actually needed, which is a word you can use.
Expect the plateau, and change tactics when you hit it. Drilling is very efficient on the steep part and very inefficient on the flat part. Somewhere around 1,500 to 2,000 characters, reading a lot of slightly-too-hard material starts outperforming any amount of flashcard discipline, because the remaining characters are rare enough that you only really learn them by meeting them in context.
How we counted
The figures on this page come from Hanzio's own frequency data. It is proprietary, built in-house from large corpora of Chinese in both spoken and written registers, weighted so that no single register dominates.
It comes to 33 billion Han characters and 19,361 distinct characters. Each corpus was counted separately and normalised over its own total before the weights were applied, so a big corpus cannot drown a small one. Characters were counted as they occur in running text, including inside multi-character words, which is why the totals are character counts rather than word counts. Every percentage on this page is a share of character tokens - of the characters you actually meet - not a share of the dictionary.
Three limits worth stating plainly. This is simplified Chinese only - we have not measured traditional, so none of these numbers should be quoted for it. The blend is a model, not a measurement of your reading: it describes a reader who splits their attention across spoken and written Chinese in those proportions, and nobody does exactly that. The corpora drop the rarest items, which nudges the coverage figures very slightly upward - the effect is small, but it runs in the optimistic direction rather than the pessimistic one, and you should know which way it leans.
Hanzio shows a frequency figure on every dictionary entry: as a word, as a building block inside longer words, and as a component of other characters. Those are scores on a 0-100 scale rather than the rank used here, and they keep word and building-block use separate where this page adds the two together. So they are broadly correlated with the ranking used here rather than identical to it: 不 comes fifth on this page's curve, but first on the building-block measure and last of that group on the standalone-word one. Same data, different question.
Frequently asked questions
Is 2,000 characters enough to read a Chinese newspaper?
At 2,000 characters you know about 97% of the characters in running text, which sounds close but works out at roughly one unknown character every 34 - several per paragraph. You can follow a news article with a dictionary open and guess a fair amount from context. Comfortable, uninterrupted reading needs more than characters though: it needs the vocabulary they spell out, so treat 3,000 as the character half of the job rather than the whole of it.
How many characters do I need for 90% of everyday Chinese?
About 1,090. The curve is steep at the start and flat at the end: 500 characters already covers 75% of running text and 1,000 covers 89%, but each further 500 adds less than the 500 before it. Getting from 90% to 95% costs roughly another 500 characters.
How many Chinese words do I need to know?
Far more than characters. Reaching the same 89% coverage that 1,000 characters gives you takes around 11,800 words. That is the practical case for learning characters early - they are the smaller and more reusable unit. It is not a case for learning them instead of words, because knowing two characters does not tell you what their combination means.
How many characters does a native speaker know?
Educated adult native speakers are usually estimated at around 6,000, and Chinese schooling targets roughly 3,500. The full national standard table lists 8,105. Those are inventory counts of what a standard contains, which is a different question from how much running text a given number of characters covers.
Do these numbers apply to traditional characters?
No. Everything here was counted on simplified Chinese, so the figures describe mainland material. The shape of the curve - very steep at the start, very flat at the end - holds for traditional characters as well, but we have not measured the traditional numbers, so we are not going to quote any.
Which characters should I learn first?
The most common ones, and the order matters far more at the start than later on. The thirty most frequent characters alone account for about 23% of everything you read. Any frequency-ordered list will do the job; Hanzio shows a frequency figure on every dictionary entry so you can see where a character sits before deciding to study it.
Do these figures include names and rare characters?
Mostly. The counts come from our own frequency lists, built over raw text, so personal names, place names and brand names are all in there, which is deliberate - they are part of what you actually meet on a page. The lists do drop their rarest entries, though, so the real tail is a little longer than these figures show: the last 1% of coverage is spread across thousands of characters you will meet once, and there are somewhat more of them than we counted.
Do I need to be able to write the characters by hand?
Not for reading. Recognising a character and producing it from memory are different skills, and every figure on this page is about recognition only. Handwriting does help many learners tell lookalike characters apart, and it is how most people finally notice that 未 and 末 are not the same shape, but it is a separate goal.
Learn them in the order that pays
The curve only helps if you know where a character sits on it. Hanzio puts a frequency figure on every dictionary entry, so when you look something up you can see at a glance whether it is common enough to be worth your time now, or something to leave for later. The dictionary is free.