IDP Glossary: Intelligent Document Processing Terms Explained
Production IDP has its own vocabulary. Some terms are borrowed from adjacent fields and used loosely. Others are used precisely in one context and differently in another.
This glossary covers the terms that matter most in production document extraction pipelines — defined from two years of running live systems, not from vendor documentation.
Each entry explains what the term means, how it works in practice, and where it matters most. Where a concept warrants more depth, there’s a link to a full guide.
All terms
- Confidence Scoring in Document Extraction: What It Is and Why It Matters
- Human-in-the-Loop Document Processing: What It Is and How to Design It
- Schema-First Extraction: What It Is and Why It Matters for Production IDP
- Structured vs Unstructured Documents: What's the Difference?
- What is a Document Extraction Pipeline?
- What is Document Automation?
- What is Document Classification in IDP?
- What is Document Validation in Extraction Pipelines?
- What is Layout Variation in Document Extraction?
- What is OCR (Optical Character Recognition)?
- What is OCR Post-Processing?
- What is Straight-Through Processing (STP)?
- What is Table Extraction from PDFs?