Document Character Counter
Drop documents and get precise character counts: characters with and without spaces, letters, digits, lines, punctuation and the UTF-8 file size of the text, for Word, PDF, OpenDocument, RTF, HTML and text files. Characters are counted the way people see them, so emoji and accented letters count once, and the Unicode code point and UTF-16 counts used by programs are shown too.
- Runs in your browser
- No sign-up
- Free to use
| Document | Characters | Without spaces | Letters | Digits | Lines | UTF-8 size |
|---|
Extracted text (to check what was counted)
How to use Document Character Counter
- Drop one or more documents.
- Read characters with and without spaces per file.
- Check the totals, bytes and Unicode counts.
- For short texts, compare with SMS and post limits.
Document Character Counter features
Human-correct counting
Grapheme clusters: an emoji or é counts as one character.
Every count
With and without spaces, letters, digits, punctuation, lines.
Programmer counts
Unicode code points, UTF-16 length and UTF-8 bytes.
Many formats
DOCX, PDF, ODT, RTF, HTML, Markdown and plain text.
Limits
X post, SMS segments (GSM-7 or Unicode), titles and descriptions.
Private
Files are read in your browser.
When to use Document Character Counter
- Checking a translation or abstract against a character limit.
- Billing translation work per character.
- Seeing how many SMS segments a message will need.
- Checking text length against a database column’s byte limit.
Document Character Counter FAQ
Do spaces count as characters?
Both are shown. Many limits, such as those of journals and translation agencies, specify whether spaces are included.
Why do different tools give different counts for emoji?
An emoji can be several Unicode code points and even more UTF-16 units. This tool counts what users see as one character, and also shows the code point and UTF-16 counts that programs use.
What is the UTF-8 size?
The number of bytes the text takes when stored as UTF-8. English letters take one byte, accented letters two, most Asian characters three and emoji four.
How are SMS segments calculated?
Texts using only the GSM-7 alphabet fit 160 characters in one SMS (153 per part when split); any other character switches to Unicode, with 70 (67) per part.
Can it read scanned PDFs?
Only PDFs with a text layer. Use the PDF OCR tool for scanned documents.
Are my files uploaded?
No. Everything happens in your browser.
What counts as a character
Character limits appear everywhere: abstracts, form fields, advertisements, translation quotes, SMS messages and database columns. Counting is harder than it seems because the word “character” means different things to people and to computers. A family emoji looks like one symbol but is made of several Unicode code points, and JavaScript would report an even larger length.
This counter reports what users see by default. It uses the browser’s grapheme segmentation, so a letter with an accent, a flag or a combined emoji counts as one character. Alongside, it shows Unicode code points, which most programming languages count, UTF-16 code units, which JavaScript and Java report as length, and UTF-8 bytes, which determine storage size and many database limits.
Breakdowns help with specific rules. Characters without spaces are common in academic and translation contexts; letters and digits separate text from numbers; punctuation and spaces explain the difference between counts; line counts matter for subtitles and code. Non-ASCII characters are counted because they increase the byte size and change how some systems store the text.
For short texts, practical limits are shown. SMS messages are split into segments whose size depends on whether every character exists in the GSM-7 alphabet; a single character outside it, such as many emoji or letters with uncommon accents, reduces the capacity from 160 to 70 characters per message. Posts, page titles and meta descriptions have their own usual lengths.
Text is extracted from documents in your browser: Word and OpenDocument files are unpacked, PDFs read with PDF.js, RTF and HTML reduced to their text. The extracted text is shown for checking, and nothing is uploaded.