Scans and photos become data,
without a person retyping them
Paper and scanned documents still arrive by email, courier and messenger photo, and someone has to read them and type the fields into a system. We build an extraction agent on a vision model that reads the document, pulls the fields you need, and flags anything it is not confident about instead of guessing.
The process today
Plenty of a business’s real documents never arrive as clean digital data: a passport photo sent on a messenger, a scanned ID, a supplier’s handwritten delivery note, a faxed form that nobody has retired yet. Someone has to look at each one, read the relevant fields, and type them into whatever system needs the data, and that someone usually gets it right most of the time but tires late in a long batch, exactly when the error rate climbs.
The second cost is turnaround. A customer or applicant waiting for a document to be processed is waiting on whoever has time that day to read it, which can mean hours or days of delay on something a machine could read in seconds if it were set up to.
The third is volume spikes. A process that works fine at ten documents a day falls over at two hundred, because the bottleneck is a person’s reading speed, not the complexity of the task.
What the agent does
The agent reads scans, photos and PDFs with a vision model built for this kind of work, which handles rotated pages, uneven lighting and handwriting far better than traditional OCR. It extracts the specific fields your process needs, mapped to your schema, whether that is a passport number and expiry date, an invoice total, or a delivery note’s item list.
Every extracted field comes with a confidence score. Fields above your threshold flow straight into your database, spreadsheet or system of record; fields below it go to a human review queue with the original image attached and the uncertain field highlighted, so the reviewer spends seconds, not minutes, checking just the part that needs a second look.
Where it is useful, the agent also validates extracted data against records you already hold, catching a passport number that does not match the name on file, or a delivery note that does not match an open purchase order, before the data gets used downstream.
What stays with humans
Anything the agent is not confident about is a human decision, not a guessed value saved to your system. Legal or compliance judgment calls about a document’s authenticity, and any edge case the extraction schema was not built for, get routed to a person rather than forced through the pipeline. The schema itself, what fields matter and how strict the confidence threshold should be, is set by your team.
Guards
Every document processed is logged with the extracted fields, the confidence scores, and whether a human corrected anything, which both improves the system over time and gives you an audit trail. The confidence threshold starts conservative at launch, meaning more documents go to human review than will eventually be needed, and is tightened only after we can show accuracy on your actual documents, not a generic benchmark.
Price and timeline
| Option | Price | What it covers | Timeline |
|---|---|---|---|
| Single automation | from $500 | One document type, one schema, confidence routing | 3 to 8 days |
| Department package | from $2,500 | OCR and extraction plus form and application processing and data migration across your operations team | 2 to 4 weeks |
Running cost is usually $10 to $60 a month in model usage depending on document volume, with a budget cap set before launch.
Related
Pair this with form and application processing when the extracted document is part of a larger application, and with data migration between systems when extracted records need to land in a new system cleanly. For invoices specifically, see invoice processing. The full package breakdown is on the AI agents service page and the automation-everything overview; for a real build of this kind of extraction, see the visa appointment and passport OCR case study and the visa centre support bots case study.
Ready to stop retyping what a scan already says? Get in touch and we will look at a sample batch of your documents in the first call.
Tired of doing this by hand? We can take the whole routine off your team, not just this step: Routine takeover, from $400 →
FAQ
How much does document OCR automation cost?
from $500 for one document type and schema; more document types or languages add build time, quoted after a sample batch.
How long does it take to set up?
3 to 8 days, including a test run on a batch of your real documents.
What does it connect to?
Your database, spreadsheet, CRM or document management system, plus the vision model handling the actual reading.
What if the AI misreads a document?
Low-confidence reads are routed to a human instead of being guessed and saved; the threshold is set conservatively at launch and can be tightened once accuracy is proven on your documents.
Where do our documents go?
Documents are processed through the vision model's API under your account where possible, and originals stay in your storage, not ours.