← Glossary
Definition

Document Classification

Document Classification is the automatic identification of a document's type, such as an invoice, contract, identity document, or compliance filing, based on its content or appearance, so the right extraction and validation workflow can be applied to it.

Before any field can be extracted from a document, something has to decide what kind of document it is. A human reviewer does this instinctively in the first few seconds of opening a file. An automated pipeline has to do it explicitly, and getting it wrong sends a document down the wrong extraction path entirely.

Why it is the first step, not an afterthought

In a real document-heavy workflow, an inbound stack might contain invoices, purchase orders, identity documents, and compliance filings all mixed together, arriving by email, upload, or scan, with no consistent naming or structure. Sorting that stack into the right categories before anyone reads the content is classification work, and it is one of the things machine learning genuinely does reliably at volume, which is why it sits at the front of most serious IDP pipelines rather than being handled manually.

Discuss this with AI:
Last updated 9 September 2026