PDF OCR
Recognition is slow by nature, please wait for long documents
How to Use
📖 Tool Introduction
A tool for running character recognition on scanned PDFs. Upload, choose a language, and each page is recognised in turn, with the resulting text ready to copy and edit.
This is a different job from PDF to Text and applies to a completely different kind of file. That tool reads the text layer a PDF already carries and does nothing for scans; this one renders each page into an image first and hands it to a recognition engine, so it exists for scans, photographed documents and PDFs made from screenshots. Choosing between them is easy: if you can select the text with your mouse, use PDF to Text; if you cannot, use this.
The recognition engine runs entirely on your own device and never sends the file to a server, which is also why it feels slower than online services: every calculation lands on your CPU. Expect a few seconds to a dozen per page, and be patient with long documents; the progress bar shows the current page and recognition percentage, and you can stop midway, with pages already finished kept.
Languages cover Simplified Chinese, English and a mix of both. The render scale decides how much each page is enlarged before recognition, with 1.5x the sensible default; especially small type or blurry originals justify 2x for better accuracy at the cost of speed. Both the Chinese and English recognition models ship with the plugin, so there is nothing to download and it works offline once installed.
The Chinese model already contains Latin letter and digit shapes, so choosing the mixed option recognises English words and numbers inside a Chinese document, though proper nouns come out slightly less accurate than with the dedicated English model. If a document is mostly English, pick English for speed and accuracy. Conversely keep the Chinese model for predominantly Chinese text and do not switch just because a few English words appear.
✨ Key Features
- Built for scans: render then recognise, filling the gap the text extractor leaves
- Recognition engine runs fully locally, no file ever uploaded
- Simplified Chinese, English and a mixed mode
- Adjustable render scale, raise it for blurry or small-type documents
- Live progress with the option to stop, keeping finished pages
- Copy results in one click or download as txt
🚀 How to Use
- Click the upload area to pick a scanned PDF, or drag one in
- Choose the language, using the mixed option for documents combining both
- Pick a render scale, the 1.5x default suits most cases
- Press start and wait while pages are recognised one by one
- Review the result, then copy it or download and save
⚠️ Things to Note
- Slow recognition is normal: All computation lands on your local CPU, seconds per page, so be patient with long files
- Results need proofreading: Handwriting, decorative fonts and very low resolution drop accuracy noticeably
- Language data ships locally: Chinese and English models are bundled with the plugin, usable with no network at all
- Layout is not restored: Output is a plain stream of text, so tables and columns need rearranging yourself
- Very large files hit memory limits: Hundreds of pages take a very long time, so recognise in batches