Invoices and forms read into structured fields
Extraction with a confidence threshold and a human review queue for anything below it.
Documents arrive as PDFs and scans in a dozen layouts. Somebody reads each one and types the numbers into a system, and the error rate is invisible until it is a reconciliation problem.
How we would approach it
Ingest PDFs and scans, OCR where required, normalise page structure
Extract to a schema with per-field confidence rather than a blob of text
Validate against business rules — totals, tax, master data, duplicates
Route anything below threshold to a review queue with the source page beside it
Learn from corrections so the threshold can be tightened on evidence
What we would measure
Agreed at the start, reported against honestly. If a measure is not worth arguing about now, it is not worth reporting later.
- Straight-through rate
- Field-level accuracy on a held-out sample
- Review queue depth and ageing
- Corrections by field, to find the weak template
Capabilities involved
Straight-through rate is the wrong target on its own. A system that processes everything and gets ten percent subtly wrong is worse than one that escalates a quarter of it.
Have something worth building?
Tell us the constraint you are working against. If we are not the right people for it, we will say so.
Or write to connect@jannex.in