DOCUMENT DATA EXTRACTION & DOCUMENT INTELLIGENCE

Reading the text was
never the hard part.

Knowing what the document is, whether it is the current version, whether the table continues on page four, whether the totals reconcile, and which single field nobody should trust — that is the hard part. It is also where document projects quietly go wrong for months without erroring once.

Invoices, tax forms, statements, contracts, field notes, spreadsheets, presentations.Including the ones photographed at an angle, which is usually the real test.

THE PIPELINE BENCHPICK A DOCUMENT TYPE
2,400 invoices a month · 60 suppliers · no two templates alikeThe easy-sounding one. Every vendor demo uses invoices.
  1. Scanpassed
  2. Classifypassed
  3. Extractpassed
  4. Validatebreaks here
  5. Structurepassed
  6. Indexpassed
  7. Routepassed

Where it actually breaks

Validation, not extraction. The line items were read correctly and they do not add up to the stated total, because the supplier applied a credit on a separate page that no one told the system about.

What the pipeline does about it

Totals are checked against line items, tax is recomputed rather than trusted, and a mismatch becomes a flagged exception with both numbers shown — never a silent overwrite of one with the other.

Into the ledger with a confidence score per field. Twelve of 2,400 reach a human, and each one arrives with the reason attached.

THE INPUTS.
WHAT ARRIVES.
Scanned PDFsPhotosTablesHandwritingSpreadsheetsEmail

01 / WHY THE DEMO ALWAYS WORKS

Five things that break
every accuracy claim.

Vendor benchmarks are run on documents that have none of these. Your documents have several. That gap is why a 99% demo becomes a 70% reality on your files.

01

Layout complexity

Merged cells, nested tables, multi-column pages, footnotes that look like values. The text extracts fine; the relationships between it do not.

02

Handwriting

Field notes, signatures, a figure written in a margin. Present in more business documents than anyone plans for, and absent from most test sets.

03

Cross-page references

A table continuing onto page four, a total that lives three pages from its line items, a credit note that changes an invoice elsewhere in the batch.

04

Scan quality variation

The same form arrives clean from one client, photographed at an angle from another, and faxed at 150 dpi from a third. One pipeline has to survive all three.

05

Domain conventions

A field that means one thing in your industry and something else everywhere else. No general model knows your conventions, and no benchmark tests them.

This is not our opinion. Independent 2026 benchmarking of document extraction makes the point directly: benchmark scores predict little about behaviour on documents with exactly these five characteristics, because standard evaluation sets exclude them. Which is why we ask for your worst documents before quoting anything, rather than your cleanest.

02 / WHAT WE ACTUALLY BUILD

Twelve stages.
OCR is one of them.

Most tools sell you stage four and leave you the other eleven. The eleven are where the cost, the errors and the eventual trust in the output all come from.

Not every project needs all twelve, and we will tell you which ones yours does not. But the order matters: indexing documents for AI before they have been classified and validated produces a system that answers fluently from the wrong version of a contract, which is worse than having no system at all.

03 / THE ENGINE QUESTION

LLM or OCR is
the wrong question.

They fail in opposite directions, which is exactly why production pipelines in 2026 run both and route between them.

CLASSICAL OCR & DOCUMENT AI

Textract, Google Document AI, Tesseract

  • Faster and materially cheaper at volume
  • Deterministic — the same input gives the same output, every time
  • Strong on clean, consistent, high-volume forms
  • Degrades badly on poor scans, merged cells and handwriting
  • Will not invent a value it cannot read, which is a genuine virtue
MULTIMODAL MODELS

Gemini, GPT-4o, Claude and similar

  • Roughly 10–15 accuracy points better on degraded inputs
  • Handles complex tables, handwriting and unusual layouts
  • Understands context, so it can tell a footnote from a figure
  • Non-deterministic, slower and more expensive per page
  • Will produce a plausible value where it should have refused

So we route by document, not by preference. Clean high-volume forms go the cheap deterministic route; awkward, degraded and unusual documents go to the model. Where both run and disagree, that disagreement is itself the signal — it is one of the most reliable ways we have found to catch a wrong value before it reaches your ledger.

04 / THE PART THAT MAKES IT TRUSTED

Being wrong quietly
is the real failure.

A pipeline that errors is annoying. A pipeline that produces confident, plausible, wrong numbers for eight months is expensive — and it is the far more common outcome, because nothing alerts anyone.

Score the field.
Not the document.
  1. Confidence is per field, not per document

    A document-level score hides the one value that matters. Scoring each field means a review queue that contains twelve fields rather than four hundred documents, and a reviewer who knows exactly where to look.

  2. Thresholds are yours, and they differ by field

    A supplier name being slightly wrong is an annoyance. A bank account number being slightly wrong is a fraud incident. The same threshold for both is how pipelines cause damage, so we set them per field with you.

  3. Provenance travels with the value

    Every extracted figure keeps which document, which version, which page and when it was read. When someone challenges a number nine months later, that question has an answer rather than an investigation.

  4. The system is allowed to refuse

    An illegible page produces a specific re-request the same day. A classifier that is unsure holds the file rather than guessing it into the wrong pile. Refusing is cheaper than being confidently wrong.

  5. Corrections train the next run

    What a reviewer fixes is captured, so the same failure does not arrive every month forever. The measure of a good pipeline is that the review queue shrinks.

05 / THE ENGAGEMENT

Send the documents
you are embarrassed by.

The clean ones tell us nothing. The photographed, faxed, half-legible, wrongly-filed ones decide the architecture — and they are the ones a vendor demo will never be run on.

  • 01A read of a real sample of your documents, including the bad ones
  • 02Document classification, deduplication and version ranking at intake
  • 03Legibility and completeness assessment before extraction
  • 04Layout-aware extraction covering tables, multi-column pages and handwriting where present
  • 05Validation rules built from your business logic, not generic checks
  • 06Normalisation of dates, currencies, units, names and identifiers
  • 07Per-field confidence scoring with thresholds agreed field by field
  • 08A review queue that shows the source page beside the extracted value
  • 09Indexing for retrieval with citations, so the output is genuinely AI-ready
  • 10Routing into your ledger, CRM, database or downstream workflow, with provenance
What sits outside the scope
  • Guaranteeing a headline accuracy percentage before seeing your documents. Anyone quoting one is quoting a benchmark, not your files.
  • Data-entry outsourcing. We build the system; we are not a keying vendor.
  • Rekeying your historic archive beyond an agreed scope. Backfiles are their own project with their own economics.
  • Deciding what your extracted data means. We deliver correct structured data; the business judgement stays yours.

RELEVANT WORK

At volume, in production.

Site photographs and PDFs into 50-page client reports, 1,200+ times a year

A US property-services firm produces more than twelve hundred reports a year from thousands of files per engagement — photographs, PDFs, spreadsheets and field notes. The system classifies, extracts, models and assembles the report; an analyst reviews the exceptions and owns the final judgement.

Every design decision described on this page came out of work like that rather than from a whitepaper: per-field confidence exists because document-level scores hid the one value that mattered, and provenance exists because someone eventually asks where a number came from.

Read the report-production case study

06 / START WITH A SAMPLE

We will tell you
where it would break.

Send a representative batch — deliberately including the bad ones. You get a straight read on per-field accuracy against your own documents, which stages would need building, and whether this is worth automating at your volume.

  1. What your documents actually do to a pipeline
  2. Per-field accuracy on your files, not a benchmark
  3. A scope and a fixed price, in writing

You will be talking to the people who would build it, not an account manager.info@chronexa.io

Rather write it down first?

Tell us what the documents are, roughly how many, and where the output has to go.

A FEW GOOD QUESTIONS

Before you start.

Is this just OCR with a nicer wrapper?

No, and the distinction is the entire point of this page. OCR turns pixels into characters, which is the part that has largely been solved. What has not been solved is turning a document into data you can rely on: working out what the document is, whether it is a duplicate or a superseded draft, whether the table on page four continues from page three, whether the totals actually reconcile, which fields are uncertain enough to need a person, and how to keep provenance so a number can be defended a year later. In our experience the extraction step is rarely where a document pipeline fails. It fails at classification, validation or structure — all of which sit either side of OCR.

Should we use an LLM or traditional OCR?

Most production pipelines in 2026 run both, and that is what we build. Multimodal models — Gemini, GPT-4o, Claude and similar — measurably outperform classical OCR on degraded scans, complex tables and handwriting, by roughly ten to fifteen accuracy points on difficult inputs. Classical engines such as AWS Textract, Google Document AI and Tesseract remain better on clean, consistent, high-volume documents where speed, cost and deterministic output matter more than handling the awkward cases. Routing each document to the right engine, and reconciling when they disagree, is a large part of the engineering. Anyone telling you it is one or the other is selling the one they have.

What accuracy can you guarantee?

Not a number, and you should be wary of anyone who offers one before seeing your files. Published benchmarks predict remarkably little about real documents, because the characteristics that break extraction in practice — layout complexity, handwriting, references across pages, varying scan quality and domain-specific conventions — are precisely what standard evaluation sets avoid. What we will commit to is measured on your documents: we run a sample, report per-field accuracy on your actual files, and agree thresholds from that. A pipeline that is 99% accurate on the fields that do not matter and 90% on the one that does is a bad pipeline with a good headline.

Our documents are not in English, or not only in English.

That is normal and it is worth correcting a common piece of shorthand: mainstream OCR engines have supported a hundred-plus languages for years, so language alone is rarely the blocker. What actually causes trouble is the combination — a regional script in an unusual layout, at poor scan quality, with domain conventions no general model has seen, and often mixed languages within one document. We handle those together rather than treating language as a checkbox, and where a document is translated we keep the original alongside so a reviewer can always see what the source really said.

What does "AI-ready" actually mean?

It means the documents can be retrieved from and cited, not just stored. In practice: classified and organised into a schema that reflects your business, chunked in a way that respects the document's own structure rather than cutting tables in half, indexed so a query returns the right passage, and carrying enough provenance that an answer can point to a specific page of a specific version. Most "we put our documents into AI" projects fail at exactly this layer — the files were uploaded, but nothing was structured, so the AI answers confidently from the wrong version of a contract.

How much human review is there, and does it ever go away?

At the start, more than people expect — and that is deliberate, because early review is how thresholds get calibrated on your real documents rather than on an assumption. It then falls as validation rules tighten and corrections feed back. It does not go to zero, and we would be suspicious of a system that claimed it had: some documents genuinely are ambiguous, and the right behaviour is to surface them. The number we care about is how much of the queue is genuine ambiguity rather than the same avoidable failure arriving every month.

What does an engagement look like?

It starts with your documents, not a workshop. Send a representative sample — deliberately including the worst ones, because those decide the architecture — and we will tell you where a pipeline would break, what per-field accuracy looks like on your actual files, and which parts are worth automating. That produces a scope, a definition of what counts as working, and a fixed price before any build starts.