Guides
How to get the text out of a PDF
Extract selectable text as plain text or Markdown — and what to do when a PDF is really just scanned images.
There are two very different kinds of PDF, and they behave completely differently when you try to copy the words out.
PDFs with real text
Most PDFs exported from a word processor or browser store the actual text. You can pull it out as a plain-text file, or as Markdown — which infers headings, lists, and paragraphs from the layout so the result drops neatly into a wiki, README, or notes app.
PDFs that are just pictures
A scanned document is a stack of images with no text behind them. Extraction returns little or nothing because there are no words to read — only pixels. This is expected, not a bug.
Reading text from scanned images (OCR) is on the roadmap; for now, text extraction works on PDFs that already contain text.
Plain text vs. Markdown
- Plain text: a clean, linear stream — best for search, notes, or feeding another tool.
- Markdown: keeps structure (headings and lists), with pages separated by a horizontal rule.
Both run in your browser, so even a confidential document stays on your device.