Scanned PDFs are just images — no text layer, so search returns nothing and copy-paste doesn't work. OCR (Optical Character Recognition) analyzes the images and adds an invisible text layer beneath, making the document searchable and copy-pasteable in every PDF viewer. The Rankato OCR tool uses Tesseract.js — an open-source engine that runs entirely in your browser. First-run downloads ~15 MB of language models (cached afterward); processing takes ~2-5 seconds per page.
Language support
English, Spanish, French, German, Italian, Portuguese, Dutch, Russian, Chinese (simplified + traditional), Japanese, Korean, Hindi, Arabic — Tesseract handles 100+ languages. Pick the language matching your PDF for best accuracy (mixed-language docs can use multiple).
Accuracy expectations
Clean, high-resolution scans (300 DPI, no skew): 95-99% accuracy. Older or lower-quality scans: 85-95%. Handwriting is not reliably supported. Highly stylized fonts, tables, and multi-column layouts often need cleanup. Preview the extracted text before finalizing.
Processing time
Per-page time depends on resolution and page complexity — expect 2-5 seconds/page on a modern laptop. First run downloads language models (~15 MB per language), which is cached. Large scanned batches can take minutes; the tool shows progress.