Documents and AI
Why it sometimes reads wrong
What makes automatic reading fail, and what you can change.
Automatic reading does not guess: it recognises. When it fails it is almost always because there was nothing recognisable in the original, and that is something you can fix before uploading.
The five causes, by frequency
| Cause | What you see | Fix |
|---|---|---|
| Crooked or shadowed photo | Half-read or invented values | Retake it with the paper flat on a table, in good light |
| Low-resolution scan | Confuses 8 with B, 0 with O | Rescan at 300 dpi |
| Handwritten document | Very low confidence | Manual review: that is expected |
| Unusual layout | Picks up some fields, not others | Correct once; it improves for the next ones |
| Several documents in one PDF | Mixes data from two | Split them before uploading |
The confidence indicator
Every extracted value carries a note of how sure it is. High means you can use it without looking; medium, worth a glance; low, must be checked. It is not decoration: it is the difference between reviewing everything and reviewing 10%.
Important
Never use a low-confidence value for anything with consequences — a payment, an official registration — without looking at it. Automatic reading saves time, not responsibility.
›Does it learn from my corrections?
Yes, for that document type within your organisation.
›Can I turn reading off?
Yes. The document is stored just the same, only without extracted data.
›How long does it take?
Seconds for a normal document; a little longer for a hundred-page one.
A real case
The situation
Tax IDs from one supplier's certificates come out wrong again and again.
What you do
- Looks at the original: they are hand-held photos with shadow
- Asks for scans instead
What you get
From then on confidence is high and they stop being checked one by one.
The situation
The document is scanned at low quality.
What you do
- Rescans at a better resolution
What you get
Reading improves without touching anything else.
The situation
The text sits on a patterned background.
What you do
- Reviews by hand whatever is flagged
What you get
You correct the little that fails.
The situation
The document's text is an image inside a PDF.
What you do
- Checks it was processed as an image
What you get
The text becomes searchable.
The situation
A stamp covers part of a figure.
What you do
- Corrects the figure by checking the image
What you get
The figure ends up right even though the paper made it hard.
The situation
One document type always fails the same way.
What you do
- Checks whether the model fits that format
What you get
The problem is tackled at source.
This article answers
- the extracted data is wrong
- why is the tax id read incorrectly
- ocr is not reading my document
- improve reading accuracy