Working with scans

Why can't I select the text in my PDF?

Because the PDF is almost certainly a scan — a photograph of a page rather than text — so there is nothing to select until OCR reads the pixels and turns them into characters.

5 min read

You open a PDF, try to drag across a sentence, and the whole page turns blue in one block. Or nothing highlights at all. Search finds no matches for a word you can plainly see on the screen.

Nothing is broken. The document simply does not contain the text you are looking at.

Two kinds of PDF wearing the same coat

The PDF format can hold two very different things, and the file extension does not tell you which you have.

A text PDF stores characters as characters. It was produced by software — exported from Word, printed to PDF from a browser, generated by a billing system. Each glyph is recorded with its font and position. Select a line and you select those characters.

A scanned PDF stores a picture. Somebody put paper through a scanner or photographed it, and the result is a bitmap wrapped in a PDF container. The page *looks* like text because it is a photograph of text, but as far as the file is concerned it is one rectangular image. There is nothing to select but the image.

The three-second test: try to drag-select a line. If the highlight follows the words, it is text. If the whole page selects as one object, it is a scan.

A second test that catches an important middle case: search for a word you can see. If search finds nothing, there is no text layer — even if selection seemed to do something.

The middle case: searchable scans

Some scanned PDFs have already been through OCR, which adds an invisible text layer positioned behind the image. You see the scan; the search index sees the recognised text.

These behave like text PDFs for searching and copying, though the copied text sometimes contains recognition errors that the eye does not notice on the image. If you copy from a scan and get l where you expected 1, that is why.

What OCR does, and what it does not

Optical character recognition looks at groups of pixels and works out which letter each group most likely represents.

What you get back is the words. What you do not get back is the design. OCR recovers content, not layout. A recognised page gives you the text in reading order; it does not give you the columns, the table structure or the exact positions, because it is inferring those from a picture too.

This is the most common misunderstanding about OCR, and it causes real disappointment. People expect a scanned invoice to come back as a formatted invoice. It comes back as the words that were on the invoice.

PDF to Text does this for scanned PDFs, and Image to Text does it for photographs.

What affects accuracy

OCR quality varies enormously with the input, and the differences are worth knowing because most of them are controllable.

Resolution. 300 DPI is the practical standard. At 150 the letters have too few pixels and similar shapes get confused. Above 600 there is no benefit and files get large.

Straightness. A page scanned at a slight angle is much harder. Most scanners can deskew automatically; it is worth turning on.

Contrast. Clean black on white is ideal. Grey text, coloured backgrounds and photocopies of photocopies all reduce accuracy.

The typeface. Ordinary printed serif and sans-serif faces are recognised very well. Decorative fonts, condensed type and italics are harder. Handwriting is a different problem entirely and general-purpose OCR does poorly on it.

Language. Recognition uses knowledge of the language to resolve ambiguity. Telling it the right language makes a real difference, and it matters a great deal for non-Latin scripts.

Rough expectations: a clean 300 DPI scan of printed text gets into the high nineties per cent. A phone photo at an angle in poor light might reach eighty, which sounds acceptable until you realise it means one wrong character in every five words.

If you just need to fill it in or sign it

A large share of "I cannot select this text" is not really about the text. Someone has a scanned form and needs to complete it, or a contract and needs to sign it.

For that, OCR is the wrong tool and converting is unnecessary. What you want is to type on top of the image — the scan stays as it is, and your text goes over it. The OCR PDF is built for exactly that case: add text, add a signature, save a real PDF, with the original scan untouched underneath.

This is almost always faster and produces a better-looking result than any recognise-and-rebuild route.

Converting a scan to Word

If you send a scanned PDF to a PDF-to-Word converter, an honest one gives you a Word file containing a picture of each page. That is the truthful conversion: there was no text, so there is no text.

The alternative — running OCR and building a Word document from the recognised words — gives you editable text with approximate layout and whatever recognition errors came along. Sometimes that is what you want. Often it is more work to fix than retyping the parts you actually needed.

Decide by asking what you will do with it. Editing every paragraph? OCR. Changing one number and signing? Edit the image directly.

Making a scan searchable without changing it

There is a middle path worth knowing. Rather than converting a scan into another format, you can add an invisible OCR text layer to the PDF itself. The document looks exactly the same, prints exactly the same, and becomes searchable and copyable.

For archives — anything you might need to find later — this is usually the right treatment. You lose nothing and gain the ability to search a filing cabinet's worth of paper.

Getting better scans in the first place

If you control the scanning, a few habits pay off repeatedly:

  • Scan at 300 DPI, in greyscale rather than colour for plain documents.
  • Use the flatbed rather than the feeder for anything creased or stapled.
  • Turn on deskew and despeckle if your scanner offers them.
  • For phone photos: flat surface, even light, camera parallel to the page, no shadow from your own hand. A dedicated scanning app that corrects perspective is far better than the plain camera.

Common questions

How do I tell whether my PDF is a scan?

Try to drag-select a line of text. If the highlight follows the words it is a text PDF; if the whole page selects as one block it is a scan. Searching for a visible word is a second check.

Will OCR give me back my formatting?

No. OCR recovers the words, not the design. You get the text in reading order, not the columns, tables and exact positions.

How accurate is OCR?

A clean 300 DPI scan of printed text reaches the high nineties per cent. A phone photo at an angle in poor light might reach eighty. Handwriting is much worse.

I only need to sign a scanned form. Do I need OCR?

No. Type and sign directly on top of the image with the OCR PDF tool — faster, and it keeps the original page exactly as it is.