Loading...
Loading...
Run text recognition on scanned image PDFs and extract all readable text.
Your file is deleted from our server right after processing.
Recognition works on the first 30 pages per run
Your file is uploaded securely to run this tool, then deleted from our server immediately after processing — never stored, logged, or shared.
1. Drag-and-drop or select your file in the workspace above.
2. Adjust any options this tool shows for your document.
3. Click "Process OCR PDF" to run the conversion.
4. Download the finished file — your upload is deleted from our server right after processing.
OCR PDF renders each page of a scanned PDF to an image, runs Tesseract text recognition on it, and returns the recognized text as a single plain-text file — one ocr-result.txt transcript covering every page it processed, with each page's text labeled and separated, rather than a new searchable PDF.
This step only matters for pages that are pictures of text in the first place — a scan, a fax, or a photographed document — since that's the only kind of page without real text to read directly. A PDF exported from Word, a browser, or any app that already has selectable text doesn't need OCR at all; extracting that existing text directly will always be more accurate than re-reading a rendered image of the same page. Recognition is also capped at the first 30 pages of any upload, a limit DocShift applies to keep processing time reasonable, so a longer scanned document needs to run in batches — extract the page range you need first, or split the file into 30-page chunks before running each part through separately.
A language dropdown matches recognition to the document: English by default, plus Spanish, French, German, Italian, Portuguese, Hindi, Russian, Japanese, Simplified Chinese, Arabic, and Korean. Recognition is most reliable on clean, printed text in the language you select — handwriting, low-resolution scans, and unusual fonts all come back with more mistakes no matter which language you pick. Starting from raw photos instead of a PDF? Run Scan to PDF first to build the file, then run this tool on the result.
Last updated: July 27, 2026
DocShift deletes your file from our server the moment processing finishes — nothing is stored, logged, or shared. There's nothing to install and no sign-up: your document is ready in seconds.
No — it returns a plain-text (.txt) file containing the text Tesseract recognized on each page, not a new PDF with a text layer added. Copy the transcript into whatever document format you need next.
Only ones where the pages are pictures of text — scans, faxes, or photographed documents. If your PDF already has selectable, copyable text, it does not need OCR; that existing text is already more accurate than anything recognition would produce from an image of the same page.
Yes. Recognition only processes the first 30 pages of any upload, a limit DocShift applies to keep processing time reasonable. For a longer scanned document, split it into smaller files first and run each through separately.
The Document Language dropdown covers English (default), Spanish, French, German, Italian, Portuguese, Hindi, Russian, Japanese, Simplified Chinese, Arabic, and Korean. Recognition is most accurate on clear, printed text in the language you select.
Not directly — this tool takes a PDF. Turn your photos into one first with Scan to PDF, then run OCR on the result.