Document intelligence

Procys: Document intelligence

Production LLM agents for document processing and identity verification

Context

Invoices arrive with different layouts, field names and levels of scan quality. Extracting text is only part of the work. Header fields, line items and tax amounts need to form a usable record, and uncertain values need a review path before they reach the next business system.

System and approach

The pipeline combines OCR with LLM reasoning about document layout, then applies deterministic validation to the extracted fields. Field confidence and checks against existing records help separate accepted data from exceptions. Reviewers receive the document context needed to resolve uncertain fields.

  • Extract structured header fields, line items, and tax breakdowns from varied invoice layouts.
  • Combine OCR with LLM layout reasoning before deterministic validation.
  • Attach field confidence and compare identities with existing records.
  • Route uncertain fields into exception review with context attached.

Delivered scope

Document extraction, validation and review are connected in one workflow. The engineering scope includes structured invoice data, counterparty checks and exception handling.

Technology

  • Python
  • FastAPI
  • OpenAI
  • AWS Textract
  • PostgreSQL
  • Docker