How TechTalent Built an AI-Assisted Document Verification Solution

Complex document workflows can require specialists to review large amounts of information, compare details across multiple documents and apply detailed rules before reaching a decision. When much of this work is performed manually, improving the process involves more than digitizing the documents themselves. The information still needs to be extracted, structured and checked in a […]

scroll for more

Complex document workflows can require specialists to review large amounts of information, compare details across multiple documents and apply detailed rules before reaching a decision. When much of this work is performed manually, improving the process involves more than digitizing the documents themselves. The information still needs to be extracted, structured and checked in a way that supports the people responsible for the final analysis.

This was the challenge behind a recent project successfully delivered by TechTalent. Over three months, a six-person engineering team developed an AI-assisted application for a complex document verification workflow. The solution combined document processing, deterministic validation and semantic AI capabilities, with a clear objective: make the review process more efficient while keeping specialists and existing human controls at the centre of decision-making.

The Challenge - Supporting a Complex Manual Process

The existing workflow required specialists to manually review document packages against detailed requirements and established rules. Information had to be checked within individual documents and compared across the complete package, making each case a multi-stage review rather than a simple document-reading exercise.

At the beginning of the project, the documented manual processing time averaged approximately 75 minutes for one category of case and 80 minutes for another. Final validation also involved a second specialist as part of the organization’s existing control process.

The variety of documents added another layer of complexity. The project scope covered 11 commonly used document types, including commercial invoices, certificates, packing lists, transport documents and insurance documentation. Documents could also be bilingual, combining English with Romanian or another language.

Against this background, the role of the new application was clearly defined. It would assist specialists by processing the documentation, identifying potential discrepancies and presenting the relevant information for investigation, while leaving the established review and final validation process intact.

Delivering with Six People in Three Months

That focused scope was particularly important because the project was delivered by six people over three months. Maintaining velocity required the team to concentrate engineering effort on the parts of the workflow where the application could provide practical support, without attempting to replace the wider operational environment around it.

The resulting user journey was deliberately straightforward. A specialist uploads a combined PDF containing the relevant documentation and starts the verification process. The application then processes the documents and returns a structured view of potential discrepancies and supporting information. Previous processing runs can also be accessed through the application's history.

Behind that relatively simple experience sits a considerably more complex backend. The frontend is focused on document upload, processing progress, results, history and access to generated reports, while the document analysis itself takes place behind the interface.

Keeping these responsibilities separate helped create a simple interaction for users while allowing the engineering team to concentrate on the more demanding processing workflow underneath.

Building an Efficient Document Processing Pipeline

Once a document package is uploaded, the application needs to transform a collection of largely unstructured information into data that can be checked systematically. This required a processing pipeline capable of handling several different tasks while working towards the project's performance requirements.

The technical specification established an estimated processing target of approximately five minutes per case, with several operations designed to run in parallel. This was a project target rather than a documented final processing result.

The pipeline was designed to:

  • separate the combined PDF into individual documents;
  • perform OCR and extract document structure;
  • classify each document;
  • extract relevant information into structured fields;
  • parse the source information required for verification;
  • apply deterministic validation rules;
  • perform semantic checks where interpretation is required;
  • identify and classify potential discrepancies;
  • generate an annotated PDF and structured result.

Several stages, including OCR, classification, structured extraction and semantic verification, were designed to run in parallel where appropriate.

Once the documents had been converted into structured information, the next challenge was deciding how that information should be verified. This is where the architecture deliberately used different approaches for different types of checks.

Combining Deterministic Rules With AI

Some validations are based on clearly defined information and do not require semantic interpretation. The application's deterministic rules engine was designed to handle checks involving dates, amounts and tolerances, locations, required documents, formatting requirements and values that need to remain consistent across multiple documents.

Other situations are less straightforward. Comparing descriptions, checking the consistency of names and addresses or analyzing conditions expressed in less structured language can require an understanding of context rather than a direct comparison between two values.

For these checks, the architecture incorporates a semantic verification layer using an LLM together with retrieval from the relevant reference material. The system can provide the appropriate contextual information for each verification and return a structured result with a justification, while cases involving ambiguity can be marked for further human attention.

Using these approaches together allowed the application to match the verification method to the nature of the problem. Deterministic logic could handle explicit checks predictably, while AI could support areas where language and context needed to be considered.

Making the Results Practical for Specialists

The value of that processing ultimately depends on how easily a specialist can understand and investigate the findings. A list of discrepancies without clear links to the underlying documents would simply introduce another layer of manual interpretation.

For this reason, the application was designed to connect its findings directly to the source material. Potential discrepancies can be highlighted in an annotated version of the original PDF using bounding boxes generated from document coordinates. The output can also show the discrepancy type, relevant reference, expected value and the value identified in the document.

The specification defines two levels of discrepancy. Clear violations of deterministic rules or explicit conditions can be distinguished from softer cases involving ambiguity or information that requires additional attention.

Presenting the findings in this way gives specialists both the identified issue and the context required to examine it. That visibility is particularly important because the application was designed to support human review rather than make the final decision itself.

Keeping Human Expertise in the Process

Although AI is an important component of the solution, final decisions remain with specialists. The application does not independently accept or reject cases, and the organization’s existing secondary validation process remains unchanged.

Uncertainty is also made visible rather than hidden:

  • Poor-quality scans can be flagged with reduced confidence scores.
  • Handwritten documents can be detected and escalated for manual handling.
  • Unclassified or incomplete documents can be flagged for further attention or stop processing when essential information is missing.

This allows automation and AI to support the review process while preserving human judgement where it is required.

Delivering Within an Enterprise Environment

The application also needed to fit within an existing enterprise technology environment, which meant that document processing and AI capabilities were only part of the engineering challenge.

The solution combines a Chrome extension built with TypeScript and React, a .NET backend running on Azure and Azure services supporting document intelligence, AI capabilities, storage and persistence.

Enterprise requirements were incorporated into the technical design alongside the application functionality. These included integration with the organization’s identity infrastructure, encrypted document storage, protected API communication, restricted access to stored information and structured application logging.

Bringing these elements together within a three-month delivery period required the team to balance application functionality, processing requirements and the constraints of the wider enterprise environment without expanding the scope beyond the core problem the application was intended to address.

What This Project Demonstrates

The project provides a practical example of how AI can be integrated into a specialized enterprise workflow alongside conventional software engineering. Document intelligence supports extraction and layout analysis, deterministic rules handle explicit checks, and semantic AI supports areas requiring contextual interpretation. These capabilities come together in an application designed to help specialists identify and investigate relevant information while preserving the existing human review process.

It also demonstrates what a focused engineering team can deliver within a demanding timeframe when the scope and technical responsibilities are clearly defined. A six-person team delivered the project over three months, bringing together frontend and backend engineering, cloud services, document processing and AI-assisted verification around a specific operational challenge.

For TechTalent, this type of project reflects the value of assembling engineering expertise around the problem that needs to be solved. Complex software and AI initiatives often require several technical capabilities to work together, particularly when new technology must operate within existing enterprise processes and controls.

Build Your Next Engineering Team with TechTalent

TechTalent helps organizations access engineering expertise across software development, cloud, data and AI, with teams built around the technical requirements and delivery objectives of each project.

If you are planning a complex technology initiative and need the engineering capacity to move from requirements to working software, contact us to discuss your project.

Top Picks

The Benefits of Partnering with a Dedicated Development Team

The Benefits of Partnering with a Dedicated Development Team

TechTalent and SITA open a development center in Romania

TechTalent Software and SITA Partner to Open a Research and Development Center in Cluj-Napoca

press release TechTalent and Banca Transilvania tech partnership

TechTalent, a new technology partner for Banca Transilvania

How to Set Up a Dedicated Nearshore Development Center

How to Set Up a Dedicated Nearshore Development Center