Document intelligence
Kore Labs: Document mapping
A vector-first document mapping pipeline with human review for edge cases
Context
PDFs, spreadsheets and exported files needed to map into a common target schema. Their source fields and layouts varied, making a separate hand-coded mapping for every input difficult to maintain. Reviewers needed a way to inspect mappings when the system encountered an unfamiliar field.
System and approach
The mapping flow retrieves similar historical mappings through vector search before using an LLM for unknown fields. Python services coordinate the transformations, while an Angular and Node.js interface supports human review. Batch execution runs through Cloud Run and Google Cloud Storage.
- Review mappings through an Angular and Node.js interface.
- Orchestrate structured transformation with Python, FastAPI, LangChain, and Pandas.
- Retrieve similar approved mappings before model fallback for unknown fields.
- Run batches through Cloud Run and GCS while preserving human approval.
Delivered scope
The scope covers retrieval-assisted schema mapping, structured transformations, batch processing and a review interface for exceptions.
Technology
- Angular
- Node.js
- FastAPI
- LangChain
- OpenAI
- Vector Search
- Cloud Run
- GCS