Skip to content
PaperZero

Extract Text

Extract positioned PDF text into paragraphs, headings, lists, and detected tables. Choose visual or two-column reading order, remove repeated headers and footers, run optional local OCR on scanned pages, and download TXT, Markdown, or JSON with bounding boxes and font metadata.

  • Processed locally in your browser
  • No watermark
  • No sign-up
  • Works offline

Choose a PDF or drop it here

Text positions, reading order, and OCR stay on this device

How it works

  1. 1Select a PDF.
  2. 2Optionally limit extraction to specific pages.
  3. 3Choose visual or column-aware reading order and whether repeated margins should be removed.
  4. 4PDF.js geometry is grouped into lines and semantic blocks; text-free pages can use local OCR.
  5. 5Review the reconstructed text, then copy it or download TXT, Markdown, or positioned JSON.

Your privacy

This tool runs entirely inside your browser using JavaScript and Web Workers. Your document is never uploaded to any server — you can even disconnect from the internet and keep working once the page has loaded.

Frequently asked questions

My PDF is scanned and returns no text?

Enable local OCR in this tool to include scanned pages in the export. Use OCR PDF when you also need a searchable PDF copy.

Is reading order perfect for two-column papers?

No heuristic is perfect, but column-aware mode reads the left column before the right and treats centered or wide headings as dividers. Preview the result before export.

What is included in JSON?

JSON includes page dimensions, semantic blocks, table candidates, links, and every text item with its bounding box, font, size, rotation, style, and direction.

Donate