Document intelligence
Procys: Document intelligence
Production LLM agents for document processing and identity verification
Context
Invoices arrive with different layouts, field names and levels of scan quality. Extracting text is only part of the work. Header fields, line items and tax amounts need to form a usable record, and uncertain values need a review path before they reach the next business system.
System and approach
The pipeline combines OCR with LLM reasoning about document layout, then applies deterministic validation to the extracted fields. Field confidence and checks against existing records help separate accepted data from exceptions. Reviewers receive the document context needed to resolve uncertain fields.
- Extract structured header fields, line items, and tax breakdowns from varied invoice layouts.
- Combine OCR with LLM layout reasoning before deterministic validation.
- Attach field confidence and compare identities with existing records.
- Route uncertain fields into exception review with context attached.
Delivered scope
Document extraction, validation and review are connected in one workflow. The engineering scope includes structured invoice data, counterparty checks and exception handling.
Technology
- Python
- FastAPI
- OpenAI
- AWS Textract
- PostgreSQL
- Docker