Agents / Document OCR

Document OCR: turns passports and forms into structured data — without the data leaving your company

MANAGED SELF-HOSTED Runs fully on-premise

What it does

  • Reads passports, ID cards and arbitrary paper forms from a photo or a scan — bilingual documents included.
  • Returns every field as structured data ready for your system — names, dates, document numbers, checkboxes.
  • Validates what it read: date formats, document-number patterns, missing fields. When it is not sure, it leaves the field blank rather than guess — an empty cell is a signal, an invented number is a liability.
  • Copes with skewed phone photos, ticked boxes and handwritten answers on paper forms.
  • Plugs into your flow through a simple API, or an upload page for your staff.

What we'll need from you

Honest list — preparing this together is most of the project and it's why the agent answers correctly.

  • a. 30–50 sample documents of each type (anonymised is fine) to tune extraction.
  • b. The list of fields you need and the format your system expects.
  • c. Where the results should go: your CRM, a spreadsheet, a database.
  • d. For self-hosted: a machine inside your network — I'll specify it.
The economics

A person re-types about 90 documents a day. The agent reads one in seconds.

Manual data entry
≈€1.20per document
Capacity~90 documents a day
Grows with volume byhiring people
Document OCR
≈€0.01per document
Capacityseconds per document
Grows with volume bycents, or electricity on-premise
Implementation is paid once — data preparation, integrations, pilot. After that the cost grows only with usage, never with headcount.
Cloud: Claude Sonnet 5 vision, ~2.2k input and ~0.4k output tokens per page. Self-hosted: a local vision model on your hardware — electricity only. Human: ~5 minutes per document.
Managed

Processing on my EU infrastructure, with image retention configured to your policy.

Self-hosted

A local vision model inside your perimeter — built for KYC and personal data. Runs on a Mac mini M5 Pro with 64 GB (≈€3,100 one-time).

UNDER THE HOOD local vision model · image pre-processing · two-step extraction and structuring · field validation · REST API + upload page
Send it a sample document
A live demo on sample passports and forms — 30 minutes.
Request a demo