SAP Document Information Extraction, part of SAP Document AI, automates the capture of structured data from business documents such as invoices, forms, and statements. This lesson introduces how SAP Document AI extracts key fields, enriches data, supports multitenancy, and reduces manual effort—while emphasizing validation and quality considerations across machine learning, template, and generative AI extraction methods.
Describing SAP Document Information Extraction
Objective
Introduction
SAP AI Document
The SAP Document AI solution helps you process large numbers of business documents containing a wide variety of content and structures. After you upload a document file to SAP Document AI, it extracts data across different sections and layouts, regardless of how the information is organized within the document.
With SAP Document AI you can:
Process more documents efficiently with fewer errors and difficulties
Increase quality and compliance mechanisms
Reduce the time required to process a document
Allow the members of your organization to focus on more relevant tasks that are in their field of expertise
Note
Always validate information extracted using SAP Document AI before using it for critical applications. While we strive for the highest possible accuracy and quality, please note that the extraction results provided may not be entirely error-free. This limitation applies to standard and custom document types. It also applies to all available extraction methods – in other words, the SAP Document AI machine learning models, generative AI, and templates.
Note also that the quality of your extraction results depends on a wide range of factors. For more information about how to get the best out of SAP Document AI, see Best Practices (Base Edition and Premium Edition).
If you want to allow your confirmed documents to be used to improve the accuracy of SAP Document AI, you can activate the data feedback collection feature. For more information, see Create Configuration, Confirm Document, and Data Protection and Privacy.
Features
Automate Information Extraction
Automate the extraction of relevant information from business documents. The Document API takes document files as input and returns extraction results for data found across different sections and layouts.
Automate Data Enrichment
Match a business document to enrichment data records based on the information extracted from the document. The Enrichment Data API takes document files as input and returns the ID of the matching enrichment data records.
Benefit from Multitenancy Support
Use this solution in tenant-aware (multitenant) applications. Run them on a shared compute unit that can be used by multiple consumers (tenants).
Hint
SAP uses the identity and position of the document-specific fields (see Extracted Header Fields (Base Edition and Premium Edition) and Extracted Line Items (Base Edition and Premium Edition)) as a feedback signal to continuously retrain the SAP Document AI machine learning models. With this approach, SAP is able to reduce errors over time when predicting field values from documents.
This is a platform functionality reused by other applications. SAP reserves the right to reject documents submitted for retraining.
For more information, see Create Configuration, Confirm Document and Data Protection and Privacy.
Technicals
Environment: SAP Document AI is available in the following environments: Cloud Foundry environment & Kyma environment
Multitenancy Support: SAP Document AI supports multitenancy. It can be used in tenant-aware applications. For information on multitenancy support, see Run SAP Document AI in a Multitenant Application.
Summary
- Automates structured data extraction from diverse business documents via the Document API, with optional data enrichment through the Enrichment Data API.
- Higher throughput, improved quality/compliance, reduced processing time, and enables staff to focus on higher-value tasks.
- Extraction results may contain errors across ML, template, and generative methods; consult Best Practices for Base and Premium Editions.
- Optional data feedback uses confirmed documents to retrain models; SAP may reject documents submitted for retraining.
- Available on SAP BTP Cloud Foundry and Kyma environments; supports tenant-aware multitenant applications.