Concepts
OCR explained: make a scanned PDF searchable
What OCR does, when you need it, and how to add a selectable text layer to a scanned PDF — locally in your browser.
OCR — optical character recognition — reads the shapes in an image of text and works out what the letters are. It’s what turns a scan (a picture of a page) into something you can search, select, and copy.
Do you actually need OCR?
If you can already select the words in your PDF with your cursor, the text is real and you don’t need OCR — use PDF to Text or PDF to Markdown instead. OCR is for the other kind: scans and photos, where the page is just an image with no text behind it.
How the OCR tool works
Each page is rendered to an image, recognized with Tesseract, and rebuilt as a searchable PDF: the original page image with an invisible text layer sitting exactly over the words. The page looks identical, but now you can select and search the text.
The OCR engine and English language data (~11 MB) download the first time you use the tool, then are cached. Everything runs in your browser — your scans never leave your device.
Getting good results
- Clean, straight, high-contrast scans read best; skewed or blurry pages lose accuracy.
- Higher detail improves recognition but takes longer.
- It’s a best-effort transcription — proofread anything important.
- English only for now; large documents are processed up to a page limit shown in the result.
Once a PDF is searchable, you can pull the text out with PDF to Text, or convert it to Markdown for notes and docs.