Processor list

This page contains detailed information on all processors offered by Document AI. You can see a list of all processors by solution type.

All Document AI processors adhere to the Data Processing and Security Terms.

Refer to the Managing processor versions documentation for more details. Also, specific processor limits apply in addition to overall product quotas and limits.

Digitize text

Enterprise Document OCR (Optical Character Recognition)

Description

Identify and extract text in different types of documents.

This processor allows you to identify and extract text, including handwritten text, from documents in more than 200 languages. The processor also uses machine learning to perform a quality assessment of a document based on the readability of its content.

Category Digitize
Functions OCR, Quality Analysis
Release stage General availability
Access status Public
Type in API OCR_PROCESSOR
Supported languages
Full list of languages
Language Name BCP 47 Tag Script Handwriting supported
Afrikaans af Latn
Albanian sq Latn
Arabic ar Arab
Armenian hy Armn
Belarusian be Cyrl
Bangla bn Beng
Bengali bn Beng
Bulgarian bg Cyrl
Catalan ca Latn
Chinese zh Hani
Croatian hr Latn
Czech cs Latn
Danish da Latn
Dutch nl Latn
English en Latn
Estonian et Latn
Filipino fil Latn
Finnish fi Latn
French fr Latn
German de Latn
Greek el Grek
Gujarati gu Gujr
Hebrew iw Hebr
Hindi hi Deva
Hungarian hu Latn
Icelandic is Latn
Indonesian id Latn
Italian it Latn
Japanese ja Jpan
Kannada kn Knda
Khmer km Khmr
Korean ko Kore
Lao lo Laoo
Latvian lv Latn
Lithuanian lt Latn
Macedonian mk Cyrl
Malay ms Latn
Malayalam ml Mlym
Marathi mr Deva
Nepali ne Deva
Norwegian no Latn
Persian fa Arab
Polish pl Latn
Portuguese (Portugal & Brazil) pt Latn
Punjabi pa Guru
Romanian ro Latn
Russian ru Cyrl
Serbian sr Cyrl
Slovak sk Latn
Slovenian sl Latn
Spanish es Latn
Swedish sv Latn
Tagalog tl Latn
Tamil ta Taml
Telugu te Telu
Thai th Thai
Turkish tr Latn
Ukrainian uk Cyrl
Vietnamese vi Latn
Yiddish yi Hebr
Processor versions
Version ID Release Channel Release Maturity Description
pretrained-ocr-v1.2-2022-11-10 Stable GA Frozen model version of v1.0: Model files, configurations, and binaries of a version snapshot frozen in a container image for up to 18 months.
pretrained-ocr-v2.0-2023-06-02 Stable GA Production-ready model specialized for document use cases. Includes access to all OCR add-ons.
pretrained-ocr-v2.1-2024-08-07 Stable GA The main areas of improvement for v2.1 are: better printed text recognition, more precise checkbox detection and more accurate reading order.
pretrained-ocr-v2.1.1-2025-01-31 Release candidate Public Preview v2.1.1 is similar to V2.1, and is available in all regions except: US, EU, and asia-southeast1.

For more information, see Managing processor versions.

Quotas and limits
Maximum pages (online/synchronous requests): 15
Maximum pages (batch/offline/asynchronous requests): 500
Maximum pages (imageless mode online/synchronous requests): 30
Uptraining
Sample Input File Open in new window.
Sample Output Open in new window.
Supported regions
  • asia-south1
  • asia-southeast1
  • australia-southeast1
  • eu
  • europe-west2
  • europe-west3
  • northamerica-northeast1
  • us
More information Enterprise Document OCR

Extract entities from documents

Refer to Sample datasets for sample labeled and unlabeled datasets to use for training.

Custom Extractor

Description

Extract fields from documents using generative AI or custom models; fine-tune models to accurately extract data from your documents.

Category Extract
Functions OCR, Entity Extraction
Release stage General availability
Access status Public
Type in API CUSTOM_EXTRACTION_PROCESSOR
Notes
  • If using generative AI for extraction, then:

    • Only the English language is officially supported.
    • Region availability is in the US, EU, northamerica-northeast1 and asia-southeast1.

Supported languages
Full list of languages
Language Name BCP 47 Tag Script Handwriting supported
Afrikaans af Latn
Arabic ar Arab
Azerbaijani az Latn
Azerbaijani (Cyrillic) az-Cyrl Cyrl
Belarusian be Cyrl
Bulgarian bg Cyrl
Bosnian bs Latn
Catalan ca Latn
Cebuano ceb Latn
Czech cs Latn
Welsh cy Latn
Danish da Latn
German de Latn
Greek el Grek
English en Latn
Esperanto eo Latn
Spanish es Latn
Estonian et Latn
Basque eu Latn
Persian fa Arab
Finnish fi Latn
Filipino fil Latn
French fr Latn
Irish ga Latn
Galician gl Latn
Hindi hi Deva
Croatian hr Latn
Haitian Creole ht Latn
Hungarian hu Latn
Indonesian id Latn
Icelandic is Latn
Italian it Latn
Hebrew iw Hebr
Japanese ja Jpan
Javanese jv Latn
Kazakh kk Cyrl
Korean ko Kore
Kyrgyz ky Cyrl
Latin la Latn
Lithuanian lt Latn
Latvian lv Latn
Macedonian mk Cyrl
Mongolian mn Cyrl
Marathi mr Deva
Malay ms Latn
Maltese mt Latn
Nepali ne Deva
Dutch nl Latn
Norwegian no Latn
Polish pl Latn
Pashto ps Arab
Portuguese (Portugal & Brazil) pt Latn
Romanian ro Latn
Russian ru Cyrl
Russian (Petrine Orthography) ru-PETR1708 Cyrl
Sanskrit sa Deva
Slovak sk Latn
Slovenian sl Latn
Albanian sq Latn
Serbian sr Cyrl
Swedish sv Latn
Swahili sw Latn
Tagalog tl Latn
Turkish tr Latn
Ukrainian uk Cyrl
Urdu ur Arab
Uzbek uz Latn
Uzbek (Cyrillic) uz-Cyrl Cyrl
Vietnamese vi Latn
Yiddish yi Hebr
Chinese simplified zh-Hans Hani
Chinese traditional zh-Hant Hani
Zulu zu Latn
Processor versions
Version ID Release Channel Release Maturity Description
pretrained-foundation-model-v1.5-2025-05-05 Stable GA Production-ready candidate powered by Gemini 2.5 Flash LLM. Recommended for those who want to experiment with newer models.
pretrained-foundation-model-v1.5-pro-2025-06-20 Stable GA Production-ready model powered by the Gemini 2.5 Pro LLM. Supports a quota of up to 30 pages per minute for online process requests. This model has improved quality compared to v1.5, and may have a higher latency.
pretrained-foundation-model-v1.5.1-2025-08-07 Release candidate Public Preview Public preview model powered by the Gemini 2.5 Flash LLM. This model has the same features as v1.5, and has improved adaptive few-shot learning.
pretrained-foundation-model-v1.6-pro-2025-12-01 Release candidate Public Preview Preview model powered by the Gemini 3 Pro LLM.
pretrained-foundation-model-v1.6-2026-01-13 Release candidate Public Preview Preview model powered by the Gemini 3 Flash LLM.
pretrained-foundation-model-v3.5-2026-05-26 Release candidate Public Preview Preview model powered by the Gemini 3.5 Flash LLM.

For more information, see Managing processor versions.

Quotas and limits
Maximum pages (online/synchronous requests): 15
Maximum pages (batch/offline/asynchronous requests): 200
Maximum pages (imageless mode online/synchronous requests): 30
Normalized data types

You can find more information in the Enrichment & normalization, and Create dataset pages.

Full list of normalized data types
  • dateTime as STRING
  • currency as STRING
  • money as google.type.Money
  • number as FLOAT or INTEGER
Uptraining
Sample Input File Open in new window.
Sample Output Open in new window.
Supported regions
  • asia-south1
  • asia-southeast1
  • australia-southeast1
  • eu
  • europe-west2
  • europe-west3
  • northamerica-northeast1
  • us
More information Custom Extractor

Form Parser

Description

Extract general key-value pairs (entity and checkbox), tables, and generic entities from documents in addition to OCR text.

This processor applies advanced machine learning technologies to extract key-value pairs, checkboxes, and tables from documents more than 200 languages. This processor also leverages deep learning models to extract 11 generic entities that are common in various document types.

Category Extract
Functions OCR, Form Parsing, Entity Extraction
Release stage General availability
Access status Public
Type in API FORM_PARSER_PROCESSOR
Supported languages
Full list of languages
Language Name BCP 47 Tag Script Handwriting supported
Afrikaans af Latn
Albanian sq Latn
Arabic ar Arab
Belarusian be Cyrl
Catalan ca Latn
Chinese zh Hani
Croatian hr Latn
Czech cs Latn
Danish da Latn
Dutch nl Latn
English en Latn
Estonian et Latn
Filipino fil Latn
Finnish fi Latn
French fr Latn
German de Latn
Hebrew iw Hebr
Hindi hi Deva
Hungarian hu Latn
Icelandic is Latn
Indonesian id Latn
Italian it Latn
Japanese ja Jpan
Korean ko Kore
Latvian lv Latn
Lithuanian lt Latn
Macedonian mk Cyrl
Malay ms Latn
Marathi mr Deva
Nepali ne Deva
Norwegian no Latn
Persian fa Arab
Polish pl Latn
Portuguese (Portugal & Brazil) pt Latn
Romanian ro Latn
Russian ru Cyrl
Serbian sr Cyrl
Slovak sk Latn
Slovenian sl Latn
Spanish es Latn
Swedish sv Latn
Tagalog tl Latn
Turkish tr Latn
Ukrainian uk Cyrl
Vietnamese vi Latn
Yiddish yi Hebr
Processor versions
Version ID Release Channel Release Maturity Additional fields detected Description
pretrained-form-parser-v1.0-2020-09-23 Stable GA

None

Legacy version. For best quality and full feature set, use the Form Parser v2.0.
pretrained-form-parser-v2.0-2022-11-10 Stable GA
Show fields
  • email
  • phone
  • url
  • date_time
  • address
  • person
  • organization
  • quantity
  • price
  • id
  • page_number
Recommended version. Supports generic entities and includes upgraded table, KVP, and checkbox model, as well as more than 200 languages.
pretrained-form-parser-v2.1-2023-06-26 Release Candidate Public Preview

None

Public Preview version. Same model as v2.0 with native text extraction from digital PDF files enabled.

For more information, see Managing processor versions.

Quotas and limits
Maximum pages (online/synchronous requests): 15
Maximum pages (batch/offline/asynchronous requests): 100
Maximum pages (imageless mode online/synchronous requests): 30
Uptraining
Sample Input File Open in new window.
Sample Output Open in new window.
Supported regions
  • asia-south1
  • asia-southeast1
  • australia-southeast1
  • eu
  • europe-west2
  • europe-west3
  • northamerica-northeast1
  • us
More information Form Parser

Layout Parser

Description

Extracts document content elements (text, tables, and lists) and creates context-aware chunks.

Layout Parser extracts document content elements like text, tables, and lists, and creates context-aware chunks that facilitate information retrieval in generative AI and discovery applications.

Category Extract
Functions Layout Parsing, Document Chunking
Release stage General availability
Access status Public
Type in API LAYOUT_PARSER_PROCESSOR
Notes
  • This parser supports PDF, HTML, DOCX, PPTX, and XLSX/XLSM files.
Supported languages
Full list of languages
Language Name BCP 47 Tag Script Handwriting supported
Afrikaans af Latn
Albanian sq Latn
Arabic ar Arab
Armenian hy Armn
Belarusian be Cyrl
Bangla bn Beng
Bengali bn Beng
Bulgarian bg Cyrl
Catalan ca Latn
Chinese zh Hani
Croatian hr Latn
Czech cs Latn
Danish da Latn
Dutch nl Latn
English en Latn
Estonian et Latn
Filipino fil Latn
Finnish fi Latn
French fr Latn
German de Latn
Greek el Grek
Gujarati gu Gujr
Hebrew iw Hebr
Hindi hi Deva
Hungarian hu Latn
Icelandic is Latn
Indonesian id Latn
Italian it Latn
Japanese ja Jpan
Kannada kn Knda
Khmer km Khmr
Korean ko Kore
Lao lo Laoo
Latvian lv Latn
Lithuanian lt Latn
Macedonian mk Cyrl
Malay ms Latn
Malayalam ml Mlym
Marathi mr Deva
Nepali ne Deva
Norwegian no Latn
Persian fa Arab
Polish pl Latn
Portuguese (Portugal & Brazil) pt Latn
Punjabi pa Guru
Romanian ro Latn
Russian ru Cyrl
Serbian sr Cyrl
Slovak sk Latn
Slovenian sl Latn
Spanish es Latn
Swedish sv Latn
Tagalog tl Latn
Tamil ta Taml
Telugu te Telu
Thai th Thai
Turkish tr Latn
Ukrainian uk Cyrl
Vietnamese vi Latn
Yiddish yi Hebr
Processor versions
Version ID Release Channel Release Maturity Description
pretrained-layout-parser-v1.0-2024-06-03 Stable GA General availability version for document layout analysis. This is the default pre-trained processor version.
pretrained-layout-parser-v1.5-2025-08-25 Release Candidate Public Preview Preview version powered by Gemini 2.5 Flash LLM for better layout analysis on PDF files. Recommended for those who want to experiment with new versions.
pretrained-layout-parser-v1.5-pro-2025-08-25 Release Candidate Public Preview Preview version powered by Gemini 2.5 Pro LLM for better layout analysis on PDF files. v1.5-pro has higher latency than v1.5.
pretrained-layout-parser-v1.6-pro-2025-12-01 Release Candidate Public Preview Preview version powered by Gemini 3.0 Pro LLM.
pretrained-layout-parser-v1.6-2026-01-13 Release Candidate Public Preview Preview version powered by Gemini 3.0 Flash LLM.

For more information, see Managing processor versions.

Quotas and limits
Maximum pages (online/synchronous requests): 15
Maximum pages (batch/offline/asynchronous requests): 500
Maximum pages (imageless mode online/synchronous requests): 30
Uptraining
Sample Input File Open in new window.
Sample Output Open in new window.
Supported regions
  • eu
  • us
More information Layout Parser

Explore pretrained processors

Bank Statement Parser

Description

Extract from bank statements including name, account, transactions, etc.

Category Pretrained
Functions OCR, Entity Extraction
Release stage General availability
Access status Public
Type in API BANK_STATEMENT_PROCESSOR
Notes
  • If a page of a multi-page input file is the correct document type and one of the supported versions, the processor performs entity extraction on the first supported document. If the processor doesn't find any applicable documents in the input file, the processor returns an error message.
Supported languages
Language Name BCP 47 Tag Script Handwriting supported
English en Latn
Processor versions
Version ID Release Channel Release Maturity Description
pretrained-bankstatement-v1.0-2021-08-08 Stable GA
pretrained-bankstatement-v1.1-2021-08-13 Stable GA
pretrained-bankstatement-v2.0-2021-12-10 Stable GA
pretrained-bankstatement-v3.0-2022-05-16 Stable GA This version assumes that the input file contains a single bank statement. Unlike the default version, this version does not check the input file for bank statements and will not return an error if no bank statements are found.
pretrained-bankstatement-v4.0-2023-07-31 Release Candidate Public Preview
pretrained-bankstatement-v5.0-2023-12-06 Stable GA

For more information, see Managing processor versions.

Quotas and limits
Maximum pages (online/synchronous requests): 15
Maximum pages (batch/offline/asynchronous requests): 30
Maximum pages (imageless mode online/synchronous requests): 30
Fields detected in the earliest version

You can also find this information in the Field detected page.

Full list of fields
  • account_number
  • account_type
  • bank_address
  • bank_name
  • client_address
  • client_name
  • ending_balance
  • starting_balance
  • statement_date
  • statement_end_date
  • statement_start_date
  • table_item
    • table_item/transaction_deposit
    • table_item/transaction_deposit_date
    • table_item/transaction_deposit_description
    • table_item/transaction_withdrawal
    • table_item/transaction_withdrawal_date
    • table_item/transaction_withdrawal_description
Enriched fields

You can find more information in the Enrichment & normalization page.

Full list of enriched fields
  • bank_address
  • bank_name
Normalized fields

You can find more information in the Enrichment & normalization page.

Full list of normalized fields
  • ending_balance
  • starting_balance
  • statement_date
  • statement_end_date
  • statement_start_date
  • table_item/transaction_deposit
  • table_item/transaction_deposit_date
  • table_item/transaction_withdrawal
  • table_item/transaction_withdrawal_date
Uptraining
Labeling Instructions Open in new window.
Sample Input File Open in new window.
Sample Output Open in new window.
Supported regions
  • eu
  • us

W2 Parser

Description

Extract from Form W2, including employee, employer, wages, etc.

Category Pretrained
Functions OCR, Entity Extraction
Release stage General availability
Access status Public
Type in API FORM_W2_PROCESSOR
Notes
  • If a page of a multi-page input file is the correct document type and one of the supported versions, the processor performs entity extraction on the first supported document. If the processor doesn't find any applicable documents in the input file, the processor returns an error message.
Supported languages
Language Name BCP 47 Tag Script Handwriting supported
English en Latn
Supported form/versions