Document reading turns an uploaded document into fields and line items. It is built around commercial documents — invoices in particular — and does a good deal of work beyond simply recognising the characters.
What it handles for you
- Number formats. It recognises both UK style (50,000.00) and European style (50.000,00) and converts to one consistent form, working out which is in use from the document itself.
- Units. It strips noise such as "EACH" out of quantity fields and applies conversion factors where a unit such as pounds is detected.
- Field mapping. Your configured label rules are applied first, with the service's own field recognition filling any gaps.
- Awkward line items. It copes with columns all labelled the same, unit price against total price, and quantity columns that were never properly identified.
Every step is logged
Each adjustment is recorded as it is made. When an invoice extracts wrongly, that log is how you find out which rule did it — considerably faster than staring at the original document.
Practical advice
- Expect to tune it. Document extraction is never right first time across a real supplier base. Budget for a period of adjusting label rules as new layouts appear.
- Always review before posting. Put extracted financial data in front of a person before it becomes a transaction. Extraction is very good but not perfect, and the failure mode is a plausible wrong number rather than an obvious error.
- Handle "nothing found". The service can read a document and find no line items at all — a distinct outcome from an empty invoice.
- Watch mixed number formats. A document mixing conventions is most likely to mislead the format detection.
Worked Examples
- Purchase ledger: supplier invoices extracted into a draft, reviewed by accounts, then posted.
- Expenses: receipts photographed on a phone, with totals extracted and the image kept as evidence.
- Onboarding forms: scanned paper forms turned into records rather than being retyped.