All needs
Need · Read

It is a scan, and you cannot get anything out of it

When a PDF comes from a scanner or a photo, each page is only an image. Your eye reads the text; the computer sees pixels. Search finds nothing, copy-paste gives nothing. Optical character recognition adds an invisible text layer under the image, and everything becomes possible again.

  • Searching the document finds no word, even one visible on screen.
  • Selecting text selects the whole page as a photo.
  • You are retyping by hand a document already in front of you.

Poulpi’s tip

Pick the right recognition language before running it: that affects the result more than scan quality does. A French document processed as English loses every accent.

OCR — Ìfàjáde Ọ̀rọ̀

FAQ

Common questions

Will the document look different?

No. The text layer is added under the image, invisible. You see exactly the same document, but search and selection work.

Why are there mistakes in the recognised text?

Recognition depends on sharpness, contrast and page straightness. A skewed, smudged or 150-dpi scan produces the classic confusions between “l” and “1”, “rn” and “m”. Rescanning at 300 dpi changes everything.

Is handwriting recognised?

Not reliably. Recognition is trained on printed characters. On handwriting the output is too uncertain to be useful.