Legal Data Tech

Technology · OCR

Every document, automatically readable.

Many apparent AI errors are actually OCR errors. DEPLAW automatically processes every incoming document with its own OCR module optimized for the legal field — as a foundation, not as a downstream feature.

01 · Why standard OCR isn't enough

Legal documents aren't standard text.

Most law firms use the standard OCR bundled with their software — built for general office text, not for section structures, footnotes and special characters. A standard OCR optimized for a 95% recognition rate already produces around 500 misrecognized characters on a 100-page contract — enough to distort every subsequent AI analysis.

Court documents

Headers and footers, stamps, handwritten notes, varying fonts.

Contracts

Tables, exhibits, nested numbering, section structures.

Faxes & photos

Poor resolution, skewed alignment, uneven exposure.

A typical error

A decimal comma recognized as a period shifts an amount unnoticed: what should read12.500,00 € becomes12.500.00 — interpreted, depending on downstream processing, as 12,500 or as 1,250,000. Equally critical: a misrecognized digit turns a deadline on 03.10.2026 into one on 03.10.2028. Both errors look structurally unremarkable — no human flagging catches them here.

Insight: OCR for AI — the forgotten technology for legal tech? →

02 · The pipeline

Five stages to reliable text.

Every document goes through the same structured pipeline before the recognized text feeds into classification, data extraction or RAG indexing:

1

Pre-processing

Image optimization, noise filtering and skew correction for optimal recognition quality.

2

Layout detection

Automatic detection of headings, tables and formatting.

3

Text recognition

AI-based character recognition with legal dictionaries and context analysis.

4

Post-processing

Format validation and structured output in Markdown.

5

Quality assurance

Automatic plausibility checks and confidence scoring for critical elements.

03 · Accuracy

80 instead of 500 errors.

The difference between a 95% and a 99.2% recognition rate looks small, but isn't: at 10,000 characters, that means 500 versus only around 80 faulty characters. Checking against legal dictionaries — dictionaries of legal terminology, abbreviations and typical phrasing — keeps this error rate low, especially for section references, case numbers and deadlines.

ElementRisk if wrongApproach
Section symbol (§)Incorrect legal basisLegal character recognition
Dates & deadlinesFaulty deadline managementContext-based validation
Monetary amountsIncorrect valuationsNumerical plausibility check

Confidence Scoring

Every text region gets its own confidence score. Critical elements like deadlines or monetary amounts with low confidence are automatically flagged for manual review — more efficient than a full manual re-check, but just as reliable.

04 · Fully automatic

No document goes unprocessed.

Whether scan, fax, photo or email attachment: every incoming document automatically runs through the OCR pipeline and is stored as a searchable PDF in the case file — even for case files with thousands of pages, where full-text search would otherwise be impossible. Instead of a purchased third-party tool, DEPLAW runs its own OCR module for this, connected directly to classification, data extraction and RAG indexing.

ScanFaxPhotoEmail attachmentAutomatically stored as PDFOwn OCR module

Frequently asked questions

OCR, explained briefly.

Often not. An internal review shows: a significant share of apparent AI hallucinations comes from faulty text recognition, not from the limits of the language model. Misread amounts or shifted formatting lead to interpretation errors that have nothing to do with the model itself.

A standard OCR with a 95% recognition rate produces around 500 misrecognized characters on a 100-page contract. At 99.2%, out of 10,000 characters only about 80 remain — a difference that decides whether a downstream AI analysis stays usable at all.

A dictionary of legal terminology, abbreviations and typical phrasing that the OCR-recognized text is checked against — so, for example, "§ 123 Abs. 2 BGB" stays correct instead of turning into "S 123 Abs 2 BGB".

Confidence scoring rates the recognition certainty of each text region individually. Critical elements like deadlines or monetary amounts with low confidence are automatically flagged for manual review, instead of being silently carried over incorrectly.

See OCR at work on real documents.

We'll show you, using your own document types, how reliable text recognition actually is.

Book a demo