Skip to content
PaperZero

OCR PDF

Recognize English or Spanish text with Tesseract WebAssembly, skip pages that are already searchable, and add an aligned invisible text layer while preserving every original visual page. Review page confidence and export PDF, TXT, or Markdown without uploading the document.

  • Processed locally in your browser
  • No watermark
  • No sign-up
  • Works offline

Local OCR: PDF pages and recognized text stay in this browser. The selected language model downloads from this site once and is cached for later offline use.

Choose a scanned or mixed PDF

One PDF · text-rich pages are skipped automatically

How it works

  1. 1Open a scanned or mixed PDF and choose pages and a recognition language.
  2. 2PaperZero skips text-rich pages unless you explicitly force OCR.
  3. 3Selected pages are rendered and preprocessed one at a time before local recognition.
  4. 4An invisible text layer is aligned over the unchanged visual pages and verified before download.

Your privacy

This tool runs entirely inside your browser using JavaScript and Web Workers. Your document is never uploaded to any server — you can even disconnect from the internet and keep working once the page has loaded.

Frequently asked questions

Does OCR change how my PDF looks?

No. Preprocessing is used only for recognition. The searchable PDF retains each original visual page and adds non-rendering text behind it.

Why are only English and Spanish available initially?

Language models are several megabytes each. PaperZero starts with two explicitly pinned, self-hosted models instead of silently downloading a large catalog; more languages can be added after fixture validation.

Is handwriting recognition accurate?

Tesseract is designed primarily for printed text. Handwriting recognition is labeled Beta and always preserves the original handwritten page for manual review.

Donate