AI · 6 min read
OCR and document digitisation for offices still running on paper
AAKZEN TECHNOLOGIES ·
Most digitisation projects stop at the wrong point. Files get scanned, the scans go into folders named by year, and the office now has a room of paper and a drive full of images that are just as hard to search. Money spent, problem unchanged.
The scan is the cheap half. The value is in what happens after it.
The three stages, in order
**Scanning** turns paper into images. Necessary, mechanical, and the part everyone budgets for.
**Recognition** turns images into text. This is OCR, and it is where the searchability comes from. Without it, a scanned page is a picture of words, not words.
**Structuring** turns text into fields — this document is a fee receipt, for this student, dated this, for this amount. This is the stage that makes records usable rather than merely findable, and it is the stage most often left out.
A project that stops after stage one has converted a storage problem into a different storage problem. A project that reaches stage three has replaced "send someone to look" with a query.
What the technology handles well
- **Printed and typed documents.** Largely a solved problem. Accuracy on clean printed text is high enough to rely on.
- **Structured forms.** When every document has the same layout, extracting the same fields from each is very reliable.
- **Tables and ledgers**, provided the ruling is reasonably clean.
- **Both English and Hindi**, though accuracy differs and should be tested on your actual documents rather than a sample.
What it does not
- **Handwriting.** Genuinely hard, and the more so for the fast, individual handwriting of somebody filling the same form for the twentieth time that day. Treat any confident promise here with suspicion and demand a demonstration on your real files.
- **Poor scans.** Skewed pages, shadows from a phone camera, faded thermal paper, ink bleeding through from the reverse. Scan quality caps everything downstream, and no amount of processing recovers what was not captured.
- **Carbon copies and duplicate books.** Common in Indian offices, and consistently the worst input.
- **Documents with no consistent structure.** Extraction works because of pattern. Where there is no pattern, expect a human in the loop.
Before committing to any project, take fifty of your worst documents — not your best — and have them processed. The accuracy on those is the accuracy of the project.
Sequencing it sensibly
**Start with what is actively used, not with the oldest.** The instinct is to begin at the back of the room with 1998. The value is in the files people request weekly.
**Decide the index fields before scanning.** What will people search by? Name, admission number, date, document type, department. If you do not know before you start, you will scan twice.
**Keep the originals until the process is proven.** Legal retention requirements aside, destroying paper before verifying the digital copy is a decision that cannot be undone.
**Sample-check the output.** Automated extraction has an error rate. Establish what it is on your material, decide what rate is acceptable for each document type, and build in a review step for anything that matters.
**Plan the stop point.** Digitising the backlog is worth little if new paper keeps arriving. The project should end with new documents being captured digitally at the point of creation, or you will be running it again in three years.
What it is worth
The honest answer is that this is rarely justified by staff-time savings alone on the backlog. The justification is usually one of these:
- **Retrieval time.** A request that took a day now takes a minute, which changes what the office can offer.
- **Risk.** Paper burns, floods and gets misfiled. A single copy in one room is a real exposure for records with a legal retention period.
- **Audit and compliance.** Producing a complete file on request, quickly, is often the actual requirement.
- **Space.** Occasionally the room itself is worth more than the project costs.
If none of those apply strongly, it is reasonable to digitise only what is active and leave the rest on the shelf. That is a legitimate answer, and it is not the one most vendors will offer.