Text Metrics That Matter
A word count looks like the simplest statistic in writing — until submission systems disagree with your editor. This guide covers the counting conventions, the derived metrics, and what the numbers are actually for.
Updated 2026-08-06 · ~7 min read
Counting conventions: why tools disagree
Word counting has genuine ambiguity at the edges: hyphenated compounds count as one word or two depending on the convention, numbers and symbols split counters, and whitespace-only lines differ between tools. The disagreements are small — one to two percent — but they matter exactly where precision counts: assignment limits and submission cutoffs. The professional response: know which convention the receiver uses and count to it, and when in doubt, leave margin below the limit rather than arguing over a hyphenated word. Convention awareness converts mysterious discrepancies into predictable ones.
The character count's different job
Character counts serve different constraints than word counts: platform limits (posts, meta descriptions, SMS segments), legal boilerplate boxes, and typography estimates all think in characters, with and without spaces as separate figures because formatting systems budget both differently. The practical mapping: word targets govern essays and articles; character targets govern anything displayed in fixed space. Writers who track both metrics catch the failure mode where a piece meets its word limit but overflows every container it was written for.
Sentences and paragraphs: the structure metrics
Sentence count divided into word count yields average sentence length — the readability proxy that actually predicts reading effort. Paragraph statistics serve document design: average paragraph length flags walls of text (long averages) or choppy rhythm (very short ones) before readers feel them. The editing use is direct: a draft running twenty-eight-word average sentences reads dense whatever its vocabulary; one running six-word averages reads staccato. The numbers name the problem precisely, which is halfway to fixing it.
Reading time: the estimate and its honest error bars
Reading time divides word count by an assumed speed — the standard band centers around two hundred to two hundred fifty words per minute for general content. The estimate carries known distortions: technical material reads slower, familiar topics faster, and skim-friendly list content much faster still. The honest use is planning rather than promise: a blog targeting a five-minute read needs roughly a thousand to twelve hundred words, and the estimate tells you when a draft is structurally in range. Presenting reading time as precise measurement oversells an approximation.
Keyword density: the metric SEO misused
Density — a term's occurrences divided by total words — exists to flag repetition balance, and its history is a cautionary tale: early search ranking rewarded cranked density, so writers cranked it, and ranking moved on while the practice lingered. The modern reading: density surfaces accidental over-repetition (a term appearing every paragraph reads like a broken record) and coverage gaps (the topic term appears twice in a piece about it). Chasing any target percentage is optimization theater; using density as a mirror for natural usage is editing.
The editing loop: statistics as feedback, not goals
The metrics earn their place inside a revision loop: draft, measure, identify the outlier statistic, edit toward the norm of good writing in that genre. Long sentences get split; short ones occasionally merged. Bloated paragraphs get broken; thin ones consolidated. The counter tells you where the draft deviates; judgment decides whether the deviation is a flaw or a feature. Writers who treat the numbers as diagnostics improve measurably; writers who treat them as scoreboards produce padded, gamed text that reads exactly like it was padded and gamed.
Submission and limit workflows
The high-stakes use: meeting hard limits with margin. Applications, contests, and grants specify counts that automated systems enforce — a one-word overage can disqualify otherwise strong work. The workflow that prevents disaster: check the count with the receiver's convention, keep a safety margin (five to ten percent below hard limits absorbs any convention difference), and re-verify after final edits, because last-minute changes are how drafts slip over lines. The counter's job in this flow is boring and absolute: the number it shows is the number that counts.
Draft hygiene: counting the right document
A recurring failure class: measuring the file that includes front matter, references, or annotations against a limit meant for the body. Academic submissions, blog platforms, and contest systems each define the countable region differently — body-only is the common intent, but paste-everything happens constantly. The discipline: extract exactly the region the limit governs, count that, and keep the extracted version as the submission source. Ten seconds of extraction discipline prevents the disqualification no amount of writing quality recovers from.
Content planning with the numbers
Editorial planning runs on these metrics in aggregate: a content calendar targeting certain reading times needs word-count budgets per format, and series consistency means comparable lengths across installments. The planning loop: decide the format's target length, brief writers with the number, and verify at delivery. Teams that brief with explicit word budgets get consistent series; teams that brief with 'about a page' get an anthology of mismatched pieces. The counter is the verification step that keeps the plan honest at delivery.
Privacy of drafts
Unpublished text is among the most sensitive things people routinely paste into websites — drafts leak plots, strategies, and confessions. That makes processing location a genuine feature distinction for a text tool: local analysis means the manuscript exists nowhere but the writer's browser, never in a service's logs or a training dataset. For work-in-progress especially, 'where does this text go' is the evaluation question, and on-device analysis is the only answer that keeps drafts where they belong — with their author.
Counting words where the count actually matters
Word counts carry different definitions in different systems, and mismatches cause real disputes — submission systems rejecting essays that 'their' counter says are over limit. The variations: whether hyphenated compounds count as one word or two, whether numbers count at all, whether headers and footers are included, and how whitespace variants (non-breaking spaces, em spaces from design tools) affect splitting. When a limit matters, count with the same tool the checker uses, or at least know the gap: a 2,000-word essay with fifty hyphenated terms can differ by dozens of words between counters.
Character limits behave differently and bite harder, because they have no ambiguity but no mercy either. Twitter-style limits, SMS segments, and meta description lengths count characters exactly, and spaces count — a fact that surprises everyone writing their first meta description at 161 characters. For these limits, character count with spaces is the binding measure, and word count is irrelevant. SEO meta descriptions land around 150–160 characters before search results truncate them; writing to the word count there is aiming at the wrong target.
Reading time estimation deserves its own calibration. The standard 200–250 words-per-minute assumes continuous prose; technical material, code-heavy tutorials, and anything with tables reads slower, and the honest estimate discounts for it. A 1,500-word guide with three comparison tables is not a six-minute read for most people — it is closer to nine or ten. When publishing reading-time labels, bias generous rather than optimistic: a reader who finishes early feels efficient, a reader who runs late feels misled.
Common mistakes with this tool
- Counting with the wrong convention for a hard submission limit.
- Chasing keyword-density percentages instead of natural usage.
- Measuring the whole file against a body-only limit.
- Presenting reading-time estimates as precise measurements.
Frequently asked questions
How are words counted?
Whitespace-separated tokens — the convention most submission systems follow.
Why does my word processor disagree?
Edge conventions differ around hyphenated words and symbols; the gap is small but real.
How is reading time estimated?
Word count divided by an average reading rate; technical content runs slower than the estimate.
What is keyword density for?
Spotting over-repetition and coverage gaps — not hitting a target percentage.
Is my draft private?
Yes — analysis is fully local; the text never leaves your browser.
Why do different word counters give different results?
Counting rules differ: hyphenated words, numbers, punctuation-only lines, and unusual whitespace are handled inconsistently. For strict limits, use the counter the receiving system uses.
How long should a meta description be?
Around 150–160 characters including spaces — search engines truncate beyond that. The binding measure is characters, not words; word count is the wrong target for meta fields.