Scanned PDF to Text (OCR)
Extract text from a scanned PDF — one made of photographed or scanned pages, with no selectable text layer — using OCR. Runs entirely in your browser; nothing is uploaded.
Click to upload or drag and drop
Select a PDF file
How it works
Each page of your PDF is rendered as an image, then read by an open-source OCR (optical character recognition) engine that runs as WebAssembly directly in your browser tab — the same technique the Image to Text tool uses, applied page by page. Nothing is uploaded to a server. Because every page has to be rendered and recognized individually, this is slower than regular text extraction — expect it to take a few seconds per page.
Unlike the PDF to Text Converter, which only falls back to English-only OCR when a PDF has no text at all, this tool runs OCR on every page in the language you choose (English, Spanish, French, German, Hindi or Portuguese). That makes it the right choice for non-English scans and for PDFs where only some pages are images.
Getting the best results
- Higher-resolution scans (200 DPI or better) recognize far more reliably
- Straight, non-skewed pages work better than crooked or rotated scans
- Printed text is much more reliable than handwriting
- Very large PDFs (dozens of pages) will take a while — keep the tab open while it works
Frequently asked questions
How is this different from the PDF to Text Converter?
PDF to Text reads a PDF's existing text layer and only falls back to English-only OCR when the whole file has almost no text. This tool always runs OCR on every page and lets you pick the language, so use it for non-English scans or PDFs where only some pages are images. If your PDF already has selectable text, the tool will point you to the faster converter.
Is my file uploaded anywhere?
No — each page is rendered and recognized directly in your browser tab using WebAssembly. The PDF never leaves your device.
Why is this slower than regular PDF text extraction?
Every page has to be rendered as an image and individually recognized by the OCR engine, which takes a few seconds per page, unlike reading an existing text layer instantly.
Will this work well on any scan quality?
Higher-resolution, straight (non-skewed) scans of printed text recognize best. Low-resolution photos, crooked pages, or handwriting will be less reliable.