OCR GUIDE

How to OCR a scanned PDF

OCR converts pixels in scanned pages into machine-readable text.

Choose the language

Select English, Hindi, or English plus Hindi before processing. OCR quality depends strongly on language and scan quality.

Process each page

PDF Forge renders each page to an image and runs Tesseract OCR locally in the browser. The recognized text is shown as it is produced.

Export the result

The recognized text is packaged into a real DOCX document that can be edited in Microsoft Word or compatible software.

Improve accuracy

Use high-resolution scans, straighten pages, remove heavy backgrounds and avoid very small text when possible.

Built on Hatchable