Working with scans

How do I get the text out of a screenshot or a photo?

Run the image through optical character recognition, which reads the shapes of the letters and gives you back real, selectable text.

5 min read

An error message you need to search for. A page of a book. A receipt. A slide someone photographed instead of sharing. All the same problem: the words are right there and you cannot select them.

To a computer, an image of text is not text. It is a grid of coloured dots that happens to look like letters to you. Getting the words out means recognising them — which is what OCR does.

What OCR is doing

Optical character recognition works in stages, and knowing them explains every quirk of the results.

First it finds the text: which parts of the image are writing and which are background, picture or noise. Then it separates lines, then words, then individual characters. Then it matches each character shape against what it knows of the alphabet, and finally it checks the result against a dictionary — which is how "rn" gets corrected to "m" when the surrounding word makes that obvious.

Every one of those stages depends on the image being clear enough to make the right decision. That is why image quality matters far more than which tool you use.

What makes recognition accurate

Resolution. A character needs roughly 20 pixels of height to be recognised reliably. Text photographed from across a room does not have that, no matter how sharp the photo looks to you.

Contrast. Black on white is ideal. Grey on grey is hard. Text over a photograph is the hardest case there is, because the stage that separates writing from background has no clean answer.

Straightness. Every degree of rotation makes line detection harder. A page photographed at an angle, with the far edge smaller than the near one, is harder still — the letters change size across the line.

Even lighting. A shadow across half the page, or a bright reflection from a window, can make one region unreadable while the rest is perfect.

Plain type. Ordinary printed fonts recognise almost perfectly. Handwriting is a genuinely different and much harder problem. Decorative and script fonts sit in between and are unreliable.

Getting a good capture

If you are taking the photograph yourself, four habits do more for accuracy than anything else:

  1. Fill the frame with the text. Get close. Background is wasted pixels.
  2. Shoot straight down, not at an angle.
  3. Use even light. Near a window is good; direct sunlight and camera flash both create hot spots.
  4. Hold still. Motion blur destroys the letter shapes that recognition depends on.

A screenshot is better than a photograph of a screen every time — no lens, no lighting, no angle, and pixel-perfect edges.

Doing it

Image to Text takes a JPG, PNG or WebP and gives you the recognised text. It handles several languages, and telling it which language the text is in genuinely improves accuracy, because the dictionary stage has something correct to check against.

If your source is a scanned PDF rather than an image, PDF to Text is the same job on a different container.

What to check in the result

OCR does not fail loudly. It produces confident, plausible, wrong text — which is why proofreading matters more here than in most tasks.

The reliable trouble spots:

  • 0 and O, 1 and l and I, 5 and S, 8 and B. Worst in reference numbers, where the dictionary stage cannot help because the string is not a word.
  • Decimal points and thousands separators. A misread separator changes a number by a factor of a thousand and looks perfectly normal.
  • Column boundaries. Two columns of text can be read straight across, interleaving two unrelated sentences.
  • Line breaks. Hyphenated words split across lines sometimes stay hyphenated.

If you are extracting figures — an invoice, a statement, a table — check every number. If you are extracting prose, read it once. That is proportionate: a wrong word in a paragraph is obvious, a wrong digit in an amount is not.

When OCR is not the answer

The text is already selectable. Try selecting it first. If a text cursor appears, the words are real and you can copy them directly — recognition would only introduce errors into text that is already perfect.

The document exists somewhere as a file. Asking for the original is always better than recognising a picture of it. Ten seconds of asking beats proofreading.

It is handwriting. Printed-text OCR is not built for it, and the results will be poor enough to retype anyway.

Languages, and why telling it matters

Recognition uses a dictionary to resolve ambiguous characters, and a dictionary in the wrong language actively makes things worse — it will confidently "correct" a correctly-read word into a similar-looking word from the language it was told to expect.

So selecting the right language is not a formality. It is the difference between a small number of character-level errors and a text that has been fluently rewritten into nonsense.

Two specific cases:

Mixed-language documents. A page of Bengali with English technical terms is genuinely hard, because the two need different models. Recognising it twice — once per language — and taking the good parts of each is inelegant and usually gives a better result than either pass alone.

Non-Latin scripts. Accuracy varies a lot by script and by how much the letterforms depend on context. Connected scripts are harder than separated ones, and any script where a character changes shape depending on its neighbours is harder still. Expect to proofread more, not less.

What to do with the result

Recognised text arrives as plain text with no formatting, which is right — the formatting was pixels and cannot be recovered.

If the source was a document with structure worth keeping, the practical route is to recognise the text, paste it into a word processor, apply headings and lists by hand, and produce a clean PDF with Word to PDF. You end up with a document that is searchable, selectable and a fraction of the size of the scan.

That is more work than a converter promising to do it in one step. It is also the only way to get a result that is actually correct, because the structure has to come from somebody who can read the page.

Common questions

Why is my OCR result full of mistakes?

Almost always image quality rather than the tool. Recognition needs roughly 20 pixels of character height, good contrast, a straight page and even lighting. Photographed-from-a-distance text fails all four.

Does OCR work on handwriting?

Not reliably. Printed-text recognition is a different problem from handwriting recognition, and the results are usually poor enough that retyping is faster.

Which characters get misread most often?

0 and O, 1 and l and I, 5 and S, 8 and B — worst inside reference numbers, where the dictionary check cannot help because the string is not a word.

Should I enhance the image first?

Yes, if it is grey, speckled or low-contrast. Sharpening and a black-and-white conversion fix exactly what the first stage of recognition struggles with.