PDF OCR
Turn a scanned PDF into one you can search, select and copy from. Text is recognised on your device and returned as a searchable PDF and a plain-text file.
- Files stay on your device
- No sign-up
- Free to use
How to use PDF OCR
- Drop your scanned PDF onto the upload area.
- Select the language the document is written in. Choosing the right language improves accuracy considerably.
- Optionally enter a page range, then select Recognise text. Each page takes a few seconds.
- Download the searchable PDF, the text file, or both.
PDF OCR features
Searchable PDF output
The original page images are kept and an invisible text layer is added, so the document looks the same but can be searched and copied.
Plain-text export
Get all recognised text as a .txt file, with a marker for each page.
16 languages
Including English, Bengali, Hindi, Arabic, Spanish, French, German, Chinese and Japanese.
On-device recognition
The OCR engine runs in your browser with WebAssembly. Page images are not uploaded.
Page range support
Process only the pages you need to save time on long documents.
When to use PDF OCR
- Making scanned contracts and letters findable with Ctrl+F.
- Copying text out of a scanned book chapter or handout.
- Preparing a scan for conversion with PDF to Word or PDF to Excel.
- Archiving paper records so they can be searched later.
PDF OCR FAQ
How accurate is the recognition?
Clean, straight scans of printed text at 300 DPI are typically recognised with very few errors. Accuracy drops with low resolution, skewed pages, coloured backgrounds, handwriting and decorative fonts.
Does it recognise handwriting?
Not reliably. The engine is trained on printed type. Neat block capitals may partly work; cursive handwriting generally does not.
How long does it take?
Roughly two to six seconds per page on a modern laptop, longer on phones. The first run also loads the OCR engine, which takes a few seconds.
Is my document uploaded for OCR?
No. Recognition happens entirely in your browser. For languages other than English, a language data file is downloaded from a public CDN, but your pages are not sent anywhere.
My PDF already has selectable text. Do I need OCR?
No. If you can select text in the PDF, use PDF to Text or PDF to Word directly. OCR is for documents where the pages are images.
What OCR does to a scanned PDF
A scanner produces a photograph of each page. To a computer that photograph is just pixels, which is why you cannot search a scanned PDF or copy a sentence from it. Optical character recognition analyses the image, finds lines and words, and works out which characters the shapes represent.
This tool renders each page at a resolution suited to recognition, passes it to the Tesseract engine, and receives both the text and the position of every word. For the searchable PDF, that text is placed invisibly over the page image at the matching positions. The page still looks exactly like the scan, but PDF readers can now search it and select text from it.
For the best results, scan at 300 DPI in black and white or greyscale, keep pages straight, and select the correct language. Tables and multi-column layouts are recognised but the reading order of the text file may need adjusting.