[ CASE STUDY // APPLIED AI ]

AI-Enabled Document Analysis and Process Automation

A hybrid pipeline of text recognition, rules, language processing, classification and language models – delivered as a custom business application.

Hybrid AI processing pipeline

A document passes through nine consecutive stages: PDF or scan, text recognition, rule-based analysis, language processing, entity recognition, classification, language model analysis, domain logic and structured result. The result is then handed to the business application through an interface.

  1. 01 PDF / SCAN
  2. 02 OCR
  3. 03 RULES
  4. 04 NLP
  5. 05 ENTITIES
  6. 06 CLASSIFICATION
  7. 07 LLM
  8. 08 DOMAIN LOGIC
  9. 09 RESULT

APPLICATION · API

A custom business application built in 2025. The public description is deliberately abstracted: the client, document contents, reference numbers and amounts are not named.

Starting point

Processing incoming creditor documents involves a wide range of PDF files that have to be analysed, classified and assigned to the appropriate downstream workflow.

The information needed for that is largely unstructured. Layout, wording and quality vary considerably depending on the sender, and some documents exist only as scans.

The objective was an application that prepares and analyses these documents automatically and provides the relevant information in structured form for further processing.

The challenge

A purely rule-based system reaches its limits quickly when structure and wording vary. Conversely, using a language model for every single step would be slow, expensive and less reproducible.

A hybrid architecture was therefore developed. Depending on the task, the system uses conventional software logic, pattern matching, language processing, trained classification models or language models.

The processing pipeline

1. Document ingestion

PDF files are imported and prepared for further processing.

2. Text extraction and OCR

Where a document carries no usable text layer, scanned content is made accessible through optical character recognition.

3. Rule-based analysis

Recognisable structures and known patterns are processed deterministically. Such information is identified reproducibly without invoking a model for every step.

4. Language processing

Language processing techniques analyse the extracted text and identify relevant linguistic and domain-specific structures.

5. Entity recognition

Entities and characteristic information within the document are identified and converted into structured data.

6. Document classification

A trainable text classification model supports the automatic assignment of documents to defined categories.

7. Language model analysis

For tasks requiring broader semantic or contextual interpretation, large language models can be applied.

8. Structured results and downstream processing

The results of the individual stages are combined into structured information that the application or downstream systems can process further.

Hybrid AI architecture

A core architectural principle is the deliberate combination of methods. Processing therefore does not follow the pattern document → LLM → response but a controllable pipeline:

Text
PDF / Scan → OCR → Rules → NLP → Entities → Classification → LLM → Domain logic → Result

Each stage can use the technology that best fits its requirements for accuracy, predictability, speed and operational cost.

Local and external models

The application is designed to support multiple LLM backends. Alongside external interfaces, models can be operated locally through Ollama or LM Studio.

This enables several deployment strategies:

  • fully local processing
  • use of external interfaces
  • a combination of local and external models
  • model selection depending on the individual workload

For sensitive documents, local processing offers an additional option when designing the system and data protection architecture.

More than a prototype

The AI components are embedded in a complete business application. The system includes user and role management, backend services and REST interfaces, persistent data storage, administration functionality, management of training data, configurable recognition patterns, OCR configuration, configurable LLM backends and a web interface for operation and administration.

That is what turned individual techniques into an application people can actually use.

Technologies

Backend – Python, Flask, SQLAlchemy, REST, JWT Frontend – Vue.js, Vite, Bootstrap AI and language processing – spaCy, named entity recognition, text classification, machine learning, large language models Document processing – PDF processing, OCR, Tesseract, TrOCR LLM integration – Ollama, LM Studio, external interfaces

Result

The project shows how a concrete business problem becomes a complete AI-enabled software solution. AI was not treated in isolation but developed as one component of an architecture comprising document processing, machine learning, domain logic, backend services, user interface and system integration.

The approach transfers to other document-intensive processes — contracts, invoices, insurance documents or administrative correspondence.

On the underlying capability: AI-Enabled Software Solutions.

Discuss your technical starting point