Agents / Document OCR
Document OCR: turns passports and forms into structured data — without the data leaving your company
MANAGED
SELF-HOSTED
Runs fully on-premise
What it does
- Reads passports, ID cards and arbitrary paper forms from a photo or a scan — bilingual documents included.
- Returns every field as structured data ready for your system — names, dates, document numbers, checkboxes.
- Validates what it read: date formats, document-number patterns, missing fields. When it is not sure, it leaves the field blank rather than guess — an empty cell is a signal, an invented number is a liability.
- Copes with skewed phone photos, ticked boxes and handwritten answers on paper forms.
- Plugs into your flow through a simple API, or an upload page for your staff.
What we'll need from you
Honest list — preparing this together is most of the project and it's why the agent answers correctly.
- a. 30–50 sample documents of each type (anonymised is fine) to tune extraction.
- b. The list of fields you need and the format your system expects.
- c. Where the results should go: your CRM, a spreadsheet, a database.
- d. For self-hosted: a machine inside your network — I'll specify it.
The economics
A person re-types about 90 documents a day. The agent reads one in seconds.
Manual data entry
≈€1.20per document
Capacity~90 documents a day
Grows with volume byhiring people
Document OCR
≈€0.01per document
Capacityseconds per document
Grows with volume bycents, or electricity on-premise
Implementation is paid once — data preparation, integrations, pilot. After that the cost grows only with usage, never with headcount.
Cloud: Claude Sonnet 5 vision, ~2.2k input and ~0.4k output tokens per page. Self-hosted: a local vision model on your hardware — electricity only. Human: ~5 minutes per document.
Managed
Processing on my EU infrastructure, with image retention configured to your policy.
Self-hosted
A local vision model inside your perimeter — built for KYC and personal data. Runs on a Mac mini M5 Pro with 64 GB (≈€3,100 one-time).
UNDER THE HOOD
local vision model · image pre-processing · two-step extraction and structuring · field validation · REST API + upload page
Send it a sample document
A live demo on sample passports and forms — 30 minutes.