---
id: KB-DI-004
url: https://app.codecontract.io/help/documents-and-ai/why-it-sometimes-reads-wrong
idioma: en
categoria: documentos-ia
subcategoria: lectura
audiencia: usuario
nivel: intermedio
actualizado: 2026-08-13
tambienEn: [es]
relacionados: [KB-DI-002, KB-DI-005]
citadoPor: [KB-DI-003, KB-DI-005, KB-DI-009, KB-DI-011, KB-DI-013, KB-DI-015, KB-DI-018, KB-GL-009, KB-GL-010]
---

# Why it sometimes reads wrong

_What makes automatic reading fail, and what you can change._

**Responde a:** the extracted data is wrong · why is the tax id read incorrectly · ocr is not reading my document · improve reading accuracy

Automatic reading does not guess: it recognises. When it fails it is almost always because there was nothing recognisable in the original, and that is something you can fix before uploading.

**En corto**

- A text PDF reads almost perfectly; a crooked photo does not.
- The confidence indicator tells you when to look.
- Correcting a field teaches the system for that document type.

## The five causes, by frequency

| Cause | What you see | Fix |
| --- | --- | --- |
| Crooked or shadowed photo | Half-read or invented values | Retake it with the paper flat on a table, in good light |
| Low-resolution scan | Confuses 8 with B, 0 with O | Rescan at 300 dpi |
| Handwritten document | Very low confidence | Manual review: that is expected |
| Unusual layout | Picks up some fields, not others | Correct once; it improves for the next ones |
| Several documents in one PDF | Mixes data from two | Split them before uploading |

## The confidence indicator

Every extracted value carries a note of how sure it is. High means you can use it without looking; medium, worth a glance; low, must be checked. It is not decoration: it is the difference between reviewing everything and reviewing 10%.

> [!IMPORTANT]
> Never use a low-confidence value for anything with consequences — a payment, an official registration — without looking at it. Automatic reading saves time, not responsibility.

**Does it learn from my corrections?**

Yes, for that document type within your organisation.

**Can I turn reading off?**

Yes. The document is stored just the same, only without extracted data.

**How long does it take?**

Seconds for a normal document; a little longer for a hundred-page one.

## Ejemplos

**Tax IDs from one supplier's certificates come out wrong again and again.**

- Looks at the original: they are hand-held photos with shadow
- Asks for scans instead

→ From then on confidence is high and they stop being checked one by one.

**The document is scanned at low quality.**

- Rescans at a better resolution

→ Reading improves without touching anything else.

**The text sits on a patterned background.**

- Reviews by hand whatever is flagged

→ You correct the little that fails.

**The document's text is an image inside a PDF.**

- Checks it was processed as an image

→ The text becomes searchable.

**A stamp covers part of a figure.**

- Corrects the figure by checking the image

→ The figure ends up right even though the paper made it hard.

**One document type always fails the same way.**

- Checks whether the model fits that format

→ The problem is tackled at source.
