Glossary
What is OCR
Turning a picture of text into text that can be searched and read.
OCR
Optical character recognition: the technique that looks at the image of a page and recognises which letters and numbers are on it, turning it into text.
A scanned PDF is a photograph of a piece of paper: to a computer, pixels. OCR is what turns those pixels into words, and it is what then allows searching inside it or extracting a date from it.
Why it sometimes fails
| Original | How well it is recognised |
|---|---|
| Computer-generated PDF | Perfectly: the text is already inside |
| 300 dpi scan, straight | Very well |
| Phone photo, in good light | Well |
| Crooked or shadowed photo | Patchy or badly |
| Handwriting | Badly, and that is expected |
Watch out
What you cannot read at a glance, OCR cannot either. If you hesitate looking at the document, do not expect it to come out well.
Worth knowing
A PDF that already carries text does not need OCR, which is why it reads perfectly. Asking for PDFs instead of photos is the cheapest way to improve reading.
›Does OCR change my document?
No. The original is stored as is; the recognised text is kept separately.
›Does it work in other languages?
Yes, including non-Latin scripts.
›Can misreadings be corrected?
Yes, and the correction improves reading for that document type.
A real case
The situation
A company cannot find anything inside its old scans.
What you do
- Checks they were uploaded without reading
- Reprocesses them
What you get
They become searchable by content, not just by filename.
The situation
A photo of a delivery note does not allow searching by order number.
What you do
- Checks it was processed on upload
What you get
The number is found without opening the image.
The situation
A skewed, shadowed scan reads badly.
What you do
- Retakes the capture with better light and framing
What you get
Reading improves without changing anything in the system.
The situation
A handwritten document does not read well.
What you do
- Reviews by hand whatever is flagged as uncertain
What you get
You correct the little that fails instead of keying it all.
The situation
A PDF already contained text and is processed again.
What you do
- Checks whether the document was already searchable
What you get
You avoid spending on something that was not needed.
The situation
There are hundreds of unprocessed old scans.
What you do
- Processes first the ones actually consulted
What you get
Effort concentrates on what somebody will look for.
This article answers
- what is ocr
- optical character recognition
- how a scanned pdf is read
- search inside scanned documents