Guides

How to get the text out of a PDF

Extract selectable text as plain text or Markdown — and what to do when a PDF is really just scanned images.

By The Editorial Team · Published 2026-08-01

There are two very different kinds of PDF, and they behave completely differently when you try to copy the words out.

PDFs with real text

Most PDFs exported from a word processor or browser store the actual text. You can pull it out as a plain-text file, or as Markdown — which infers headings, lists, and paragraphs from the layout so the result drops neatly into a wiki, README, or notes app.

PDFs that are just pictures

A scanned document is a stack of images with no text behind them. Extraction returns little or nothing because there are no words to read — only pixels. This is expected, not a bug.

Reading text from scanned images (OCR) is on the roadmap; for now, text extraction works on PDFs that already contain text.

Plain text vs. Markdown

  • Plain text: a clean, linear stream — best for search, notes, or feeding another tool.
  • Markdown: keeps structure (headings and lists), with pages separated by a horizontal rule.

Both run in your browser, so even a confidential document stays on your device.

Related articles