1. Separate text files from scanned images.
A PDF can contain selectable text, scanned page images or a mixture of both. Optical character recognition, or OCR, converts text in images into machine-readable characters. It does not by itself establish what a value means. A number near the bottom of an invoice might be a total, a tax amount or a page reference; layout and context are needed to distinguish them.
Inspect the range of documents before choosing a method. Include rotated pages, faint scans, tables, handwritten annotations and multi-page records where these occur in the process. Keep the original file unchanged and record where it came from. Avoid assessing an extraction approach only on clean, uniform documents if the incoming material is varied.
2. Compare recognition with interpretation.
Traditional OCR is useful when the main requirement is readable text. Layout-aware tools also identify relationships such as table cells and key-value pairs. Azure AI Document Intelligence and Amazon Textract provide document-analysis capabilities. Assess the relevant document types, supported processing routes, regional arrangements and contractual terms in their documentation.
A language model can help interpret varied wording or map extracted content into a defined structure. It can also fill gaps with plausible values that were not present in the file. Require absent information to remain absent. For stable layouts, a template or dedicated extraction tool may be easier to constrain than a general-purpose model.
Choose on the basis of field accuracy, traceability and review requirements. A system that extracts more fields is not necessarily better if staff cannot locate their source. Preserve page numbers and bounding regions where the tool supports them. This lets a reviewer compare a value with the relevant area instead of searching the entire document.
3. Define a schema before extraction.
A schema describes the fields, types and rules expected in the output. For an invoice, this might include supplier name, invoice reference, currency, issue date, line items and totals. Distinguish the original text from a normalised value. Preserve leading zeros in identifiers and do not treat every numeric-looking string as a number.
UK documents can contain date formats that become ambiguous when transferred between systems. Specify how dates are represented internally and retain the source wording for review. Keep currency explicit; a symbol alone may not be enough. Where tax information appears, extract the stated values rather than asking the model to invent a treatment or determine an organisation’s tax position.
4. Use checks that can explain a failure.
Validate required fields, formats and relationships. Compare a supplier identifier with the authorised supplier record, check whether an invoice reference already exists and examine whether stated totals reconcile with the stated components. Account for rounding and document conventions rather than assuming every mismatch means the scan was read incorrectly.
A tool’s confidence score is not a guarantee or a universal probability of correctness. Its meaning depends on the provider and output type. Decide review rules using representative documents and the consequences of an error, not a borrowed threshold. A missing payment detail or unexpected change should be escalated even when the text is confidently recognised.
Give reviewers the original page, the proposed field and the validation reason. Record corrections and preserve the original extraction result separately. This creates a useful audit trail and helps distinguish recognition errors from mistakes in the schema or business rules. Do not automatically use extracted bank details to authorise a payment.
5. Keep extraction separate from acceptance.
Passing an extraction check should create a candidate record, not imply approval for every downstream action. Define the permitted destination, duplicate handling and the route for incomplete documents. Keep upload permissions narrow and scan incoming files using the organisation’s security controls. Retention should cover both original files and derived data.
Our document AI service can be scoped around a document inventory, field specification, extraction workflow and review interface. Begin with the document types, destination system and fields that matter. Use redacted descriptions for an initial enquiry; arrangements for transferring real documents should be agreed before any confidential material is shared.
Keep an empty value honest
If a field is not present, the system should record that absence and request review rather than generate a likely answer.
