This page documents production updates to Document AI. We recommend that Document AI developers periodically check this list for any new announcements.
You can see the latest product updates for all of Google Cloud on the Google Cloud page, browse and filter all release notes in the Google Cloud console, or programmatically access release notes in BigQuery.
To get the latest product updates delivered to you, add the URL of this page to your feed reader, or add the feed URL directly.
July 17, 2026
Custom extractor model
pretrained-foundation-model-v3.5-2026-05-26 powered by Gemini 3.5
Flash LLM is available in Preview.
This processor version has ML processing capabilities in the US and EU.
For more information about available models, see the custom extractor page.
June 04, 2026
Custom extractor offers document validation and correction in Preview.
This feature allows you to enhance extraction accuracy with validation rules and document data using Common Expression Language (CEL) dialect.
For more information, see CEL dialect for document validation.
May 27, 2026
Layout parser image and table annotations is in General Availability (GA).
Layout parser can identify if there are images or tables in parsed documents. When found, images and tables are annotated as a descriptive block of text with the information depicted in the image and table.
March 31, 2026
Upgrading fine tuned custom extractor processors is now available in Preview.
The feature allows you to fine tune a new processor version with a newer base version, while keeping the configurations of the previously fine-tuned processor version selected. This is available through the UI in the Deploy & use tab in the console.
This is currently supported for upgrading pretrained-foundation-model-v1.4-2025-02-05
to pretrained-foundation-model-v1.5-2025-05-05.
For more information, see training overview.
March 27, 2026
Custom splitter model
pretrained-splitter-v1.5-2025-07-14 is available in
General Availability (GA).
March 23, 2026
Custom classifier models
pretrained-classifier-v1.6-2026-03-09 and pretrained-classifier-v1.6-pro-2026-03-09
are available in Preview.
Custom splitter models
pretrained-splitter-v1.6-2026-03-09 and pretrained-splitter-v1.6-pro-2026-03-09
are available in Preview.
March 03, 2026
Custom classifier model pretrained-classifier-v1.5-2025-08-05
is available as General Availability (GA).
For more information about available models, see the custom classifier page.
February 17, 2026
Document AI legacy processors will be discontinued on June 30, 2026. To preempt the risk of service failure while using legacy processors, we recommend transitioning to more stable, higher-quality processors.
The affected versions are:
| Type | Version |
|---|---|
| Identity parsers | pretrained-us-passport-v1.0-2021-06-14pretrained-fr-driver-license-v1.0-2021-06-14 |
| Tax and finance parsers | pretrained-1099misc-v1.1-2021-12-10pretrained-1099nec-v1.0-2021-08-11pretrained-1099r-v2.0-2022-07-25pretrained-1099int-v1.1-2021-12-10pretrained-ssa1099-v1.0-2021-08-09pretrained-1099g-v1.0-2021-05-27pretrained-1099g-v1.1-2021-12-10pretrained-1120-v3.0-2022-04-26pretrained-w9-v1.0-2020-09-25pretrained-w9-v1.1-2021-12-10pretrained-w9-v1.2-2022-01-27pretrained-w9-v2.0-2022-06-23 |
| Mortage and banking parsers | pretrained-mortgage-statement-v1.0-2021-10-17
|
| Procurement | pretrained-utility-v1.1-2021-04-09pretrained-utility-v1.2-2022-12-15 |
| Splitting | pretrained-procurement-splitter-v1.1-2021-04-09pretrained-procurement-splitter-v1.2-2022-08-19pretrained-lending-document-split-v1.0-2021-12-08pretrained-lending-document-split-v2.0-2021-12-09 |
| Summary | pretrained-foundation-model-v1.0-2023-08-22 |
To ensure uninterrupted service and benefit from improved extraction quality, we recommend you migrate to the following later versions before June 30, 2026:
Enterprise Document OCR: Migrate to
pretrained-ocr-v2.1-2024-08-07.Expense parser: Migrate to
pretrained-expense-v1.3.2-2024-09-11.Custom classifier: Migrate to
pretrained-classifier-v1.5-2025-08-05.Custom splitter: Migrate to
pretrained-splitter-v1.5-2025-07-14.Invoice parser: Migrate to
pretrained-invoice-v2.0-2023-12-06.Pay slip parser: Migrate to
pretrained-paystub-v3.0-2023-12-06.Bank statement parser: Migrate to
pretrained-bankstatement-v5.0-2023-12-06.
To learn more about the migration process, refer to Manage processor versions.
If you have any questions or require assistance, contact us at Google Cloud support.
February 16, 2026
The layout parser web interface is in Preview.
It supports a document view display for processed PDF files and supports
visualizing the document's parsed JSON, block layout, and
image or table annotation data in a user-friendly interface. Bounding box support
only exists for version processor pretrained-layout-parser-v1.0-2024-06-03.
It also supports modifying the input layout config to allow for configuring the table and image annotation feature directly from the interface.
February 09, 2026
Layout parser model
pretrained-layout-parser-v1.6-2026-01-13 powered by Gemini 3
Flash LLM is available in Preview.
This processor version has ML processing capabilities in the US and EU.
For more information about available models, see the custom extractor page.
January 27, 2026
Layout parser model
pretrained-layout-parser-v1.6-pro-2025-12-01 powered by Gemini 3 Pro
LLM is available in Preview.
This processor version has ML processing capabilities in the US and EU.
For more information about available models, see the layout parser page.
Custom extractor model
pretrained-foundation-model-v1.6-pro-2025-12-01 powered by Gemini 3 Pro
LLM is available in Preview.
This processor version has ML processing capabilities in the US and EU.
For more information about available models, see the custom extractor page.
Custom extractor model
pretrained-foundation-model-v1.6-2026-01-13 powered by Gemini 3
Flash LLM is available in Preview.
This processor version has ML processing capabilities in the US and EU.
For more information about available models, see the custom extractor page.
January 12, 2026
Document AI is introducing document-level prompting for custom document processors. This feature allows you to provide an overall description of the document to inject deep business knowledge into the model, leading to improved extraction quality.
Document-level prompting offers improved accuracy by giving the model necessary context for extraction at the document level. This allows for easily supplied general information, such as geographical limitations (for example, all the address fields are located in the USA), to guide the model.
For more details, refer to the documentation on custom extractor mechanisms and document-level prompting.
December 15, 2025
A monitoring dashboard web interface is available in Preview to monitor at the project and processor level.
You can monitor a number of metrics, such as number of successfully processed
pages and sync processing latency, across fields like location, processor_type,
and processor_id over time.
For more information, see monitoring dashboard.
November 12, 2025
Automated schema extraction for custom extractor processors is in Preview.
This feature allows you to automatically extract a document schema from a test document you supply. Then, you can approve or decline the schema and edit it manually. This saves time and effort when defining the document schema for your custom processor and allows you to focus on refining the schema.
When creating a custom extractor processor, find the Generate from document option in the Get started tab of the Google Cloud console.
November 07, 2025
Gemini layout parser is in
Preview.
The Gemini layout parser gives better layout quality on table recognition,
reading order and text recognition on PDF files. You can enable the feature by
default by selecting layout parser processor version
pretrained-layout-parser-v1.4-2024-08-25, pretrained-layout-parser-v1.5-2025-08-25
or pretrained-layout-parser-v1.5-pro-2025-08-25 for your processor.
November 04, 2025
Layout parser support for DOCX, PPTX, XLSX, and XLSM file types in Document AI is in General Availability (GA). It makes content like paragraphs, tables, lists, and structural elements like headings, page headers, and footers easily accessible. It also creates context-aware chunks that facilitate information retrieval in a range of generative AI and discovery applications.
For more information, see Process documents with Layout Parser.
October 31, 2025
Custom splitter model
pretrained-splitter-v1.5-2025-07-14 with zero-shot splitting, classification
and confidence scores is available as Release Candidate
(Preview).
October 17, 2025
Layout parser lets you parse images and tables as annotations in Preview.
Layout parser can identify if there are images or tables in parsed documents. When found, images and tables are annotated as a descriptive block of text with the information depicted in the image and table.
October 06, 2025
Capacity reservation is available for Document AI in Preview. This lets you grant capacity to selected processors and maintain a steady real-time, high-volume processing flow for document processing requests.
For the necessary steps, read the Make a capacity reservation request section of "Quotas".
Custom extractor model
pretrained-foundation-model-v1.5.1-2025-08-07 with improved adaptive few-shot
learning is available as Release Candidate
(Preview).
Support for confidence scores in Custom classifier
models pretrained-foundation-model-v1.4-2025-05-16 and pretrained-classifier-v1.5-2025-08-05
is in Preview.
For best performance, use them with fine-tuned models.
September 23, 2025
Custom classifier
model pretrained-classifier-v1.5-2025-08-05
powered by Gemini 2.5 Flash is in Preview. It has ML processing available for US and EU regions, a
maximum page limit of 30 pages,
and processing requests of 120 pages per minute.
Unlike the prior custom classifier, which used classical machine learning, this version features a new platform. It accommodates:
- High accuracy immediately, based on the document classes you define.
- Few-shot learning to further improve accuracy.
- Use of descriptions when labeling for more context and insight for document classes.
- More accurate results with the same training dataset on the fine-tuned generative AI model, compared to the trained version.
- Autolabeling documents for fine-tuning and evaluation.
- Generative AI to fine-tune and heighten accuracy.
For more information on processor versions, see Managing processor versions.
September 10, 2025
Custom Extractor version pretrained-foundation-model-v1.4-2025-02-05 will no longer be accessible on February 5, 2026.
To avoid service disruptions, migrate to a later version such as
pretrained-foundation-model-v1.5-2025-05-05 or pretrained-foundation-model-v1.5-pro-2025-06-20.
To learn more about the migration process, refer to Manage processor
versions.
September 09, 2025
Document AI supports two service tiers and associated quotas: provisioned and best effort tiers.
The base is the provisioned tier quota, which provides 120 pages per minute for Gemini 2.0 and 2.5 Flash LLM and 30 pages per minute for Gemini 2.5 Pro LLM.
If you require more volume, best effort tier quota provides 120 pages per
minute for Gemini 2.0 2.5 Flash and 60 pages per minute for
Gemini 2.5 Pro. It's only used when the provisioned quota has been
exhausted. This applies to the BestEffortOnlineProcessDocumentPagesPerMinutePerProjectUS
and EU quotas and, in the console, best_effort_online_process_document_pages_us and eu.
Best effort can get up to 240 pages per minute for custom data extractor models v1.4 and v1.5 with a quota increase request (QIR). You can make a QIR by contacting your sales team representative.
There is no service level agreement (SLA) for best effort tier.
September 03, 2025
Custom extractor model pretrained-foundation-model-v1.5-pro-2025-06-20 is
available as General Availability (GA).
For more information about available models, see the custom extractor page.
August 29, 2025
Derived entity and signature detection are now supported in custom
extractor models pretrained-foundation-model-v1.4-2025-02-05
as General Availability (GA)
and in pretrained-foundation-model-v1.5-2025-05-05, as well as pretrained-foundation-model-v1.5-pro-2025-06-20
as Preview.
Signature detection lets you identify handwritten signatures by using visual cues in the document. Derived entity detection lets you deduce entities by inference without requiring the value to be explicitly present in the text. You can use this feature to deduce the country in an address, counting items in a table, or detecting if an ID is fake.
These can be enabled in the console when creating labels or by using the
DocumentSchema.EntityType
resource in the API.
For more information, read Custom extractor with derived fields and choose label attributes.
July 22, 2025
Custom extractor model
pretrained-foundation-model-v1.5-pro-2025-06-20
powered by Gemini 2.5 Pro is in Preview.
It has ML processing available for US and EU regions, a maximum page limit of
30 pages, and processing requests
of 30 pages per minute.
For more information, see Managing processor versions.
July 04, 2025
Document AI VPC service controls (VPC-SC) integration now supports identity groups.
For more information on setting up VPC-SC identity groups, read Configure identity groups and third-party identities in ingress and egress rules.
Document AI now supports Identity and Access Management (IAM) deny policies. These policies allow you to define deny rules that prevent certain principals from using certain permissions to access Google Cloud resources, regardless of the roles they're granted.
For more information, read Deny policy overview and Document AI security and compliance.
July 03, 2025
The Document AI CDE processor now supports merging the child entities
of nested entities that extend across several pages. This is supported in custom
extractor model pretrained-foundation-model-v1.5-2025-05-05.
This change is automatic in all processors.
For customers with existing v1.5 processors, to make use of this feature, you must relabel the nested entities in different pages.
To learn more about the labeling process, refer to Label documents.
June 30, 2025
Custom Extractor model pretrained-foundation-model-v1.5-2025-05-05 is in General Availability (GA) and has fine-tuning available for the US and EU.
From version v1.4 and later, we will use a new quota for online processing called Number of online process document pages per minute per processor type and model version. This quota will be enforced at a per-page and per-foundation model level. There will be no change to the batch processing quota.
These can be enabled in the console when creating labels and by using the DocumentSchema.EntityType.
For more information, read Managing processor versions.
June 19, 2025
We've increased the maximum file size for online processing requests from 20 MB to 40 MB. This applies to all types of processors.
For more information, see the Document AI limits page.
May 19, 2025
Cross-regions importing of fine-tuned models is now supported for processor versions based on Gemini 1.5 and later, such as
custom extractors
pretrained-foundation-model-v1.2-2024-05-10 and later.
For more information, see Managing processor versions.
May 05, 2025
Custom extractor model pretrained-foundation-model-v1.5-2025-04-25 powered by Gemini 2.5 Flash LLM is available as Public Preview in US regions. The custom extractor model supports a quota of up to 15 pages per minute for online process requests.
For more information about available models, see Custom extractor model versions.
April 08, 2025
Previous Custom Extractor versions pretrained-foundation-model-v1.0-2023-08-22 and pretrained-foundation-model-v1.1-2024-03-12 will be deprecated on April 9, 2025. To ensure uninterrupted service, prediction traffic to these versions, including any fine-tuned variants, will be automatically redirected to the latest version, pretrained-foundation-model-v1.4-2025-02-05.
For guidance on how to fine-tune a new version, refer to the fine tuning documentation.
April 02, 2025
All processors can now extend the Maximum page limit for online and synchronous requests up to 30 pages.
To do so, enable imageless_mode in ProcessRequest.
For Custom Extractor, you will need to first request to be allowlisted for this feature by filling out the form Allowlist Request for 30 Page limit in CDE.
March 24, 2025
As we launch Custom Extractor version pretrained-foundation-model-v1.4-2025-02-05 in GA with fine tuning (in Preview), these versions will no longer be accessible effective September 24, 2025:
pretrained-foundation-model-v1.2-2024-05-10pretrained-foundation-model-v1.3-2024-08-31
To avoid service disruptions, migrate to a later version, such as pretrained-foundation-model-v1.4-2025-02-05. To learn more about the migration process, refer to our Manage processor versions documentation.
Customers and projects can access pretrained-foundation-model-v1.2-2024-05-10 and pretrained-foundation-model-v1.3-2024-08-31 until September 24, 2025. This includes the ability to create tuning jobs and access fine-tuned processor versions.
Starting March 24, 2025:
- Newly created processor versions using
pretrained-foundation-model-v1.2-2024-05-10can only be used for batch processing. - Newly created processor versions using
pretrained-foundation-model-v1.2-2024-05-10 and pretrained-foundation-model-v1.3-2024-08-31will have a quota limit of 120 pages per minute.
This update requires planning, but if you have questions or need assistance, contact Google Cloud support.
March 19, 2025
Custom Extractor model pretrained-foundation-model-v1.4-2025-02-05 is in General Availability (GA), and has fine-tuning available in Preview for the US and EU.
From version v1.4 and later, we will use a new quota for online processing called Number of online process document pages per minute per processor_type_and_model_version. This quota will be enforced at a per-page and per-foundation model level. There will be no change to the batch processing quota.
February 14, 2025
Custom extractor model pretrained-foundation-model-v1.4-2025-02-05 powered by Gemini 2.0 Flash LLM is available as Public Preview in US and EU regions with improved accuracy. The Custom Extractor Model supports a quota of up to 120 pages per minute for online process requests.
For more information about available models, see Custom extractor model versions.
February 03, 2025
Model pretrained-ocr-v2.1-2024-08-07 has General Availability (GA) in the US and EU.
For more information about available models, see Enterprise Document OCR and Regional and multi-regional support availability.
Model pretrained-ocr-v2.1.1-2025-01-31 is available as a Release Candidate in the regions asia-south1, australia-southeast1, europe-west2, europe-west3 and northamerica-northeast1.
For more information about available models, see Enterprise Document OCR.
January 27, 2025
For processor versions pretrained-foundation-model-v1.2-2024-05-10 and pretrained-foundation-model-v1.3-2024-08-31 custom extractors, customer-managed encryption keys (CMEK) is now supported when importing fine-tuned processor versions.
For more information, see Import processor versions.
January 23, 2025
Effective January 27, 2025, new and existing processors require explicit storage.objects.get permissions to access Google Cloud Storage buckets for training dataset imports and offline/batch processing.
You will need to review your use of training dataset imports and offline/batch processing to verify that the users of these APIs have appropriate permissions to access Google Cloud Storage buckets.
Ensure that users of these APIs have been granted one of the predefined or legacy Cloud Storage roles that includes the storage.objects.get permission (such as Storage Object Viewer). You can assign these roles in the Permissions tab of the relevant Cloud Storage bucket.
We understand that this update requires planning, but we're here to support you during this process. If you have questions or need assistance, contact Google Cloud support.
December 19, 2024
Property description is now Generally Available (GA) as part of the custom extractor in both the Document AI section of the Google Cloud console and the API, with additional support for parent entities in hierarchies.
Property description allows you to provide additional context, insights, and prior knowledge for each entity to improve extraction accuracy.
December 12, 2024
You can copy processor versions of pretrained-foundation-model-v1.2-2024-05-10 and pretrained-foundation-model-v1.3-2024-08-31 between projects by following the steps in Import a processor version.
October 22, 2024
The Document AI section of the Google Cloud console now allows you to configure property descriptions as part of the Custom extractor processor-creation process.
Property description allows you to provide additional context, insights, and prior knowledge for each entity to improve extraction accuracy.
Property descriptions can be edited after schema creation. After you update the property descriptions, you will need to either call the pretrained models or create or fine-tune a new processor version for the changes to take effect.
October 01, 2024
Custom Extractor pretrained-foundation-model-v1.2-2024-05-10 and pretrained-foundation-model-v1.3-2024-08-31 are now Stable versions.
v1.2 and v1.3 now have the following features:
- Fine-tuning is now available in Public preview.
- They were internally upgraded to a higher quality model.
- The labeling system has been upgraded to use the latest version of the OCR model.
v1.2 is recommended for the best quality. v1.3 is recommended for the lowest latency.
We recommend creating a new processor and relabeling the training and evaluation documents to benefit from both the improved quality with the new processor versions of Custom Extractor (v1.2 and v1.3) and the enhanced labeling system.
September 26, 2024
The following earlier versions of Document AI Enterprise Document Optical Character Recognition (OCR) and Expense Parser will be discontinued in the United States (US) and European Union (EU) starting April 30, 2025.
Enterprise Document OCR:
pretrained-ocr-v1.0-2020-09-23pretrained-ocr-v1.1-2022-09-12
Expense Parser:
pretrained-expense-v1.2-2022-02-18pretrained-expense-v1.3-2022-07-15pretrained-expense-v1.4-2022-11-18
We upgraded the underlying vision model for these versions to an OCR module available in the US and EU.
To ensure uninterrupted service and benefit from improved extraction quality, we recommend you migrate to the following later versions before April 30, 2025:
Enterprise Document OCR (US and EU):
- Migrate to the latest version of OCR processors:
pretrained-ocr-v2.0-2023-06-02orpretrained-ocr-v2.1-2024-08-07.
Expense Parser (US and EU):
- Migrate to the later versions of Expense Parser:
pretrained-expense-v1.3.2-2024-09-11orpretrained-expense-v1.4.2-2024-09-12. - Or migrate to the latest versions of Custom Extractor:
pretrained-foundation-model-v1.2-2024-05-10orpretrained-foundation-model-v1.3-2024-08-31.
To learn more about the migration process, refer to our Manage processor versions documentation.
If you have any questions or require assistance, contact us at Google Cloud support.
Effective April 9, 2025, the following Custom Extractor versions will no longer be accessible:
pretrained-foundation-model-v1.0-2023-08-22pretrained-foundation-model-v1.1-2024-03-12
You will need to migrate to a later version to avoid any service disruptions, such as pretrained-foundation-model-v1.2-2024-05-10 and pretrained-foundation-model-v1.3-2024-08-31 for improved quality from the latest proprietary vision models and foundation models.
We understand that this update requires planning, but we're here to support you during this process. If you have questions or need assistance, contact Google Cloud support.
September 23, 2024
Models pretrained-expense-v1.3.2-2024-09-11 and pretrained-expense-v1.4.2-2024-09-12 are available as Release Candidates (RC) for Expense Parser. They are upgrades over v1.3 and v1.4 with an enhanced underlying vision model.
For more information about available models, see Expense parser processor versions.
September 20, 2024
Custom extractor now features property descriptions.
Property description allows you to provide additional context, insights, and prior knowledge for each entity to improve extraction accuracy.
Good examples of property descriptions include location information and text patterns of the property values, which help disambiguate potential sources of confusion in the document, guiding the model with rules that ensure more reliable and consistent extractions, regardless of the specific document structure or content variations.
August 23, 2024
Model pretrained-foundation-model-v1.3-2024-08-31 is available as a Release Candidate (RC) for custom extractor. Recommended for those who want the lowest latency and best speed.
For more information about available models, see Custom extractor model versions.
Model pretrained-ocr-v2.1-2024-08-07 is available as RC version of the Document AI OCR 2.1 processor. It has three key improvements:
- Better printed text recognition.
- More precise checkbox detection.
- More accurate reading order.
August 21, 2024
Date and Currency Normalization for custom extractor
With this release, the model will deduce the region information from the document and use it to disambiguate the date and currency formats in the following ways:
- This release will enable the support of region based date and currency normalization of entities with datetime and currency data types in Custom Document Extractor (CDE) Generative AI based processor versions v1.1 and v1.2.
- Currently CDE Generative AI based processor supports date and currency normalization but it defaults to US date format and USD respectively in case the values are ambiguous. In other words, if a date can be parsed in mm/dd/yyyy and dd/mm/yyyy formats, it will use mm/dd/yyyy format for normalization. Similarly if $ can be mean USD or CAD, it would default to USD.
For more information, go to the Entity Normalization page.
July 18, 2024
For custom extractor with generative AI, model pretrained-foundation-model-v1.1-2024-03-12 provides fine-tuning for US/EU in Public preview. For more information about custom extractor models, see Custom extractor model versions.
June 04, 2024
Layout Parser in Document AI is generally available. The Document AI Layout Parser transforms documents in various formats into structured representations. It makes content like paragraphs, tables, lists, and structural elements like headings, page headers, and footers easily accessible. It also creates context-aware chunks that facilitate information retrieval in a range of generative AI and discovery applications.
For more information, see Process documents with Layout Parser.
May 28, 2024
Model pretrained-foundation-model-v1.2-2024-05-10 is available for custom extractor. Recommended for using the largest supported token limits, employing the best quality in identifying entities, or experimenting with newer models.
For more information about available models, see Custom extractor model versions.
May 06, 2024
Batch processing with Layout Parser is available. For more about Layout Parser, see Process documents with Layout Parser.
Model pretrained-foundation-model-v1.1-2024-03-12 is available for custom extractor. For more information about available models, see Custom extractor model versions.
May 01, 2024
Online processing is available for Layout Parser in Document AI. The Document AI Layout Parser transforms documents in various formats into structured representations, making content like paragraphs, tables, lists, and structural elements like headings, page headers, and footers easily accessible, and creating context-aware chunks that facilitate information retrieval in a range of generative AI and discovery applications. For more information, see Process documents with Layout Parser.
April 02, 2024
Fine tuning generative AI models within the Custom Extractor is now supported in GA. For more information, see custom processors and fine tuning pricing.
February 29, 2024
The Custom Extractor with generative AI is now available in the asia-southeast1 (Singapore) regions. For more information, see Custom processors.
See the model type, generative or custom, powering a Custom Extractor processor version by getting the model type from the processorVersions API.
The Custom Extractor supports three levels of nesting so you can easily extract structured data from complex documents and tables (earnings reports, tax forms, invoices, resumes, etc.). Learn how to use three levels of nesting.
February 16, 2024
Enterprise Document OCR version 2.0, pretrained-ocr-v2.0-2023-06-02, is now Generally Available and ready for production workloads.
Please migrate OCR workloads to this new processor version.
January 09, 2024
The Custom Extractor with generative AI has General Availability and is ready for production workloads. For more information, see the Custom Extractor with generative AI or check out the demo.
- As foundation models evolve, so will versions available within the Custom Extractor. For more information, see Managing processor versions.
- Fine tuning the foundation model within the Custom Extractor is still available, in Preview. For more information, see Fine tune and train by document type.
To better support production workloads, we reduced prices for the Custom Extractor, Custom Classifier, Custom Splitter, and Form Parser. For more information, see Document AI pricing.
Developers can now specify pages Document AI should process within a document. For more information, see IndividualPageSelector within V1 API ProcessOptions.
December 20, 2023
Custom Extractor supports fine tuning (Preview) so that you can customize foundation model results for user specific documents. This feature is available in the US region. For more information, see Fine tune and train by document type.
Custom Extractor with genAI is now available in the EU and northamerica-northeast1 regions. For more information, see Custom processors.
You can now demo genAI-powered extraction results within Custom Extractor along with output from other Document AI products such as OCR, Form Parser, and ID processing.
December 07, 2023
Enterprise Document OCR version 2.0, pretrained-ocr-v2.0-2023-06-02, has an upgraded OCR engine and model improvements. This upgrade better supports high-volume workloads.
September 25, 2023
We are launching an RC version of the pretrained-invoice-v1.5-2023-09-15 invoice processor. It includes:
- Improved base-entity extraction model for documents in English.
- Line-item grouping quality improvements.
- Better support for multi-line, multi-segment entities such as addresses and line-item descriptions.
- Enforcement of occurrence type
OPTIONAL_ONCE/REQUIRED_ONCEfor properties of nested entities. - Updated OCR engine.
September 21, 2023
Launched Document AI Enterprise Document OCR v2.0 and OCR add ons in Preview.
Enterprise Document OCR launched a Release Candidate, pretrained-ocr-v2.0-2023-06-02, which includes:
- Upgraded OCR model, optimized for various document use cases.
- Visual-element detector for boxed characters, which can increase quality up to 10% for documents with text boxes.
For more details, see the documentation, including the user guide.
OCR add ons are available from the Enterprise Document OCR processor when using pretrained-ocr-v2.0-2023-06-02. These include:
- Checkbox extraction: Detects and extracts status (marked/unmarked) in the Enterprise Document OCR response.
- Math OCR: Identifies, recognizes, and extracts formulas from documents in LaTeX output format.
- Font-style detection: Identifies word-level font properties, including type, style, handwriting, weight, and color.
For more details, see the documentation.
August 25, 2023
Document AI Workbench is now powered by generative AI with two feature launches:
Document AI Workbench Summarizer is in Preview:
- The Summarizer provides summaries for documents up to 250 pages long.
- You can customize summaries based on your preferences for length (brief, moderate, comprehensive) and format (paragraph, bullet points).
- See the user guide for more information.
Document AI Workbench custom extractor is in preview:
- Custom extractor with generative AI can help extract data from documents with free-form text (e.g., contracts) and complex layouts (e.g., invoices, W2s, bills of lading).
- The pretrained processor version, which uses generative AI, can be used out of the box without any training. Post a document to the endpoint with a list of fields to get structured data.
- Customize results by confirming content in about five documents. Workbench leverages the examples to improve accuracy using few-shot prediction.
- Extract information from documents up to 200 pages long through the asynchronous API.
- To get started, create or use an existing custom extractor to leverage a processor version.
- See the how-to guide, labeling best practices, and training use cases.
- Current limitations of generative AI extraction within the custom extractor:
- Only the English language is supported.
- Region availability is currently only in the US.
- While in preview, we recommend that you only extract up to 50 entities per endpoint with generative AI.
- When uploading a sample document to define fields and preview results on the Get started page, there can be long latencies. We're working to reduce this latency.
In addition, template-based training is available in GA within the custom extractor:
- Template-based training provides accurate predictions for documents with no layout variation (such as an application form).
- Only six labeled documents are needed to train and use a template-based processor version.
- See the user guide and training use cases.
August 18, 2023
Expense processor
A new RC version pretrained-expense-v1.3.1-2023-08-11 of the Expense processor is now available in the asia-southeast1 region for Expense Parser customers.
This release includes an improved region-based normalization, which results in an average improvement of up to a 15% accuracy on normalized date and currency entities over the current stable version.
August 01, 2023
Launched the following Document AI Workbench features:
Create and train models programmatically with more public APIs, including:
DatasetSchema APIs:
UpdateDatasetSchema,GetDatasetSchema. To create schema, use theUpdateDatasetSchemaAPI (a singleton resource).Dataset APIs:
UpdateDataset,ImportDocuments,GetDocument,BatchDeleteDocuments. Until we release aListDocumentAPI, use confirmations from theImportDocumentsAPI to create a list of documents in your dataset.
Selective labeling within Custom Document Extractor (CDE) helps you prepare a diverse set of training and test documents. Import 125+ documents to a CDE dataset, then CDE recommends documents you should label based on clustering results. Selective labeling only supports new documents imported into your dataset. If you would like recommendations for documents already in your dataset, then delete and re-import them.
Quick Tables within CDE helps you train models faster by labeling a table in bulk by applying the first row pattern to the rest of the table.
Get started with a new processor by using a default storage option for your dataset. You can still configure your own Cloud Storage location using advanced options.
July 28, 2023
Launched an update to the RC release pretrained-invoice-v1.4-2022-10-21 of the invoice processor, available to all Document AI users.
This release includes the features of the Stable version, with an average improvement of 0.03 to 0.05 micro F1 (approximately 5% to 9%) on line item entities.
July 18, 2023
The following Form Parser (pretrained-form-parser-v2.0-2022-11-10) features are Generally Available (GA):
- General field extraction: You can extract 11 different types of entities from documents.
- Enhanced check-box detection.
- Internationalization support that covers more than 200 languages.
- Upgraded key-value pair (KVP) detection model.
Form parser v2.1 (pretrained-form-parser-v2.1-2023-06-26) is in Public Preview, which uses our native PDF text extraction model on PDF documents.
The Form Parser features has the following limitations:
- Check-box doesn't support radio buttons and might not reliably parse all selection marks or keyless checkboxes.
- If there is a key without a value, the model might not parse it.
- The quality of KVP parsing might be higher for Latin languages than others.
- For tables, only simple tables are supported (no support for merged cells).
July 17, 2023
The Custom Document Splitter (CDS) within Document AI Workbench is now Generally Available (GA) for production use cases to split and classify multiple documents within a single file. With this release, all Workbench processors currently offered (Custom Document Extractor, Custom Document Classifier, and Custom Document Splitter) are available in GA.
Launched the following features for CDS:
- CDS now supports up to 1,000-page documents for async/batch prediction and up to 200-page documents when importing, labeling, training, or evaluating.
- CDS model evaluation for document split and classification
- Prepare a CDS training dataset faster by bulk labeling documents at import across multiple folders.
Released the the following enhancements for CDS:
- Improved labeling and evaluation experience with the ability to review overall document splits and classifications while viewing individual pages in a side-by-side view.
- Document names are now used in error messaging to improve troubleshooting.
- Hyphens are allowed in schema names.
June 29, 2023
Identity Document AI (IDAI) photo copy detection in ID proofing (Preview)
Updated the pretrained-id-proofing-v1.1-2023-05-18 ID proofing processor for all Document AI users.
This processor includes a new output entity fraud_signals_photocopy_detection that signals if an attached image might be a photocopy. The entity can be one of the following values: POSSIBLE_PHOTOCOPY, PASS, or INCONCLUSIVE.
June 28, 2023
The document OCR native text from digital PDF feature contains the following known issues:
- For a small number of documents, word order in lines of text that are reported by native text extraction might be inaccurate.
- Invisible text that is embedded in a native PDF might be extracted.
- Japanese documents that contain currency symbols, such as Yen, might be incorrectly extracted as
/. - Apostrophe symbols might be missing in word and/or line results.
- Native text extraction might report different word and/or line results compared to image-based OCR on an identical document.
Added fixes to our doc.proto-to-vision.proto conversion tool, which facilitates migration from Vision API TextDetection to document OCR.
The following document OCR features are Generally Available (GA). Use document OCR's configurations to optimize for stability, quality, and specific response requirements.
- Intelligent document-quality analysis
- Native text from digital PDF
- Symbol-level extraction
- Language hints
April 25, 2023
Launched the following features to improve the usability of the Document AI Workbench Custom Document Extractor (CDE):
- CDE now supports an additional 42 global languages.
- CDE lets you import processor versions across projects and processors to easily manage development and production environments.
- CDE can automatically label documents in a dataset by using a deployed processor version to help you quickly prepare training data.
Document AI Workbench Custom Document Extractor (CDE) has also made the following enhancements:
- The asynchronous prediction API can now extract data from documents up to 200 pages long.
- Improved the accuracy of extracting checkboxes.
April 17, 2023
Identity Document AI (IDAI) pricing change
We are changing the price of our US-related identity document processors. The new price is on the pricing page.
March 27, 2023
For the Document AI OCR Processor (Doc OCR), you can enable document quality assessments for all processor versions instead of a specific processor version, such as pretrained-ocr-v1.1-2022-09-12. If you enable document quality assessment, Doc OCR produces a quality score that's based on the document's readability. Quality scores range from 0 to 1, where 1 is perfect quality. Quality scores are returned in the image_quality_scores field on the Page object. All detected issues are labeled as quality or defect and sorted in descending order by confidence value. To use this feature, set process_options.ocr_config.enable_image_quality_scores= true in your API request to the OCR Processor.
The Document AI OCR Processor (Doc OCR) now has the following features:
- The OCR Processor supports language hints. The OCR engine prefers your specified languages over inferred languages. To use this feature, set
process_options.ocr_config.hints.language_hintswith a list of BCP-47 language codes in your API request to the OCR Processor. - The OCR Processor supports the option to populate symbol-level data in the document response. If enabled, the field
document.pages.symbolsis populated. To use this feature, setprocess_options.ocr_config.enable_symbol=truein your API request to the OCR Processor. - A proto converter tool that converts a
Document prototo anAnnotateFileResponseproto. This conversion lets you compare the responses between the Document AI OCR processor with the Vision API, which can help you migrate to the Document AI OCR processor from Vision API with minimal downstream changes. For details, see Document AI Toolbox. - The OCR Processor supports a heuristics layout detection algorithm, which serves as an alternative to the current ML-based layout detection algorithm. You can choose the layout algorithm that best suits your needs. To use this feature, set
process_options.ocr_config.advanced_ocr_options= legacy_layoutin your API request to the OCR Processor.
February 21, 2023
This launch upgrades the lifecycle stage of the Custom Document Extractor (CDE) component of the DocAI Workbench from Public Preview to Generally Available (GA). CDE covers essential workflows for developing custom document extraction processors with end-to-end UI support:
- Data import
- Schema creation and annotation
- Processor model training
- Evaluation and troubleshooting
- Model deployment and version management
- Human-in-the-loop (HITL) integration for "last-mile" processor quality assurance
Notable new Generally Available Custom Document Extractor (CDE) features include:
- Public APIs
- Automatic schema label creation from pre-labeled documents
- Schema label data type and occurrence editable pre-training
- New DocAI Toolkit with a labeled document converter
The following features have been upgraded:
- Processor gallery
- Schema editor
- Labeling UI
- Training pipeline
- Manage versions table
January 10, 2023
The Form Parser Release Candidate version has been renamed to pretrained-form-parser-v2.0-2022-11-10. See Document AI release notes--December 12, 2022 for more information about this release.
December 22, 2022
We are launching a public preview version of the purchase order processor, pretrained-purchase-order-v1.1-2022-06-17, with the following features:
- Support for uptraining to improve, add, and remove entities in the schema.
- Support for uptraining to add support for unsupported languages.
- Improvements to overall performance.
December 19, 2022
The Document AI OCR Processor has the following new features:
The OCR Processor now supports extracting embedded text from digital PDFs in public preview. A fallback to the optical OCR model is automatically triggered to extract text in the regions when the PDF being processed contains non-digital text. To opt into this feature, set
process_options.ocr_config.enable_native_pdf_parsing=truein your API request to the OCR Processor.Added advanced versioning support to the Document AI OCR, which enables OCR users to pin to a historical model version. When enabled, OCR outputs are guaranteed to be consistent and virtually frozen, with zero behavioral drifts. To enable advanced versioning, select the release candidate version
pretrained-ocr-v1.2-2022-11-10in your Document AI console.