This page contains detailed information on all processors offered by
Document AI. You can see a list of all processors by solution type.
All Document AI processors adhere to the
Data Processing and Security Terms .
Refer to the Managing processor versions
documentation for more details. Also, specific processor limits apply in addition
to overall product quotas and limits .
Digitize text
Enterprise Document OCR (Optical Character Recognition)
Description
Identify and extract text in different types of documents.
This processor allows you to identify and extract text, including handwritten text, from documents in more than 200 languages. The processor also uses machine learning to perform a quality assessment of a document based on the readability of its content.
Category
Digitize
Functions
OCR, Quality Analysis
Release stage
General availability
Access status
Public
lock_open
Type in API
OCR_PROCESSOR
Supported languages
Full list of languages
Language Name
BCP 47 Tag
Script
Handwriting supported
Afrikaans
af
Latn
Albanian
sq
Latn
Arabic
ar
Arab
Armenian
hy
Armn
Belarusian
be
Cyrl
Bangla
bn
Beng
Bengali
bn
Beng
Bulgarian
bg
Cyrl
Catalan
ca
Latn
Chinese
zh
Hani
Croatian
hr
Latn
Czech
cs
Latn
Danish
da
Latn
Dutch
nl
Latn
English
en
Latn
Estonian
et
Latn
Filipino
fil
Latn
Finnish
fi
Latn
French
fr
Latn
German
de
Latn
Greek
el
Grek
Gujarati
gu
Gujr
Hebrew
iw
Hebr
Hindi
hi
Deva
Hungarian
hu
Latn
Icelandic
is
Latn
Indonesian
id
Latn
Italian
it
Latn
Japanese
ja
Jpan
Kannada
kn
Knda
Khmer
km
Khmr
Korean
ko
Kore
Lao
lo
Laoo
Latvian
lv
Latn
Lithuanian
lt
Latn
Macedonian
mk
Cyrl
Malay
ms
Latn
Malayalam
ml
Mlym
Marathi
mr
Deva
Nepali
ne
Deva
Norwegian
no
Latn
Persian
fa
Arab
Polish
pl
Latn
Portuguese (Portugal & Brazil)
pt
Latn
Punjabi
pa
Guru
Romanian
ro
Latn
Russian
ru
Cyrl
Serbian
sr
Cyrl
Slovak
sk
Latn
Slovenian
sl
Latn
Spanish
es
Latn
Swedish
sv
Latn
Tagalog
tl
Latn
Tamil
ta
Taml
Telugu
te
Telu
Thai
th
Thai
Turkish
tr
Latn
Ukrainian
uk
Cyrl
Vietnamese
vi
Latn
Yiddish
yi
Hebr
Processor versions
Version ID
Release Channel
Release Maturity
Description
pretrained-ocr-v1.2-2022-11-10
Stable
GA
Frozen model version of v1.0: Model files, configurations, and binaries of a version snapshot frozen in a container image for up to 18 months.
pretrained-ocr-v2.0-2023-06-02
Stable
GA
Production-ready model specialized for document use cases. Includes access to all OCR add-ons.
pretrained-ocr-v2.1-2024-08-07
Stable
GA
The main areas of improvement for v2.1 are: better printed text recognition, more precise checkbox detection and more accurate reading order.
pretrained-ocr-v2.1.1-2025-01-31
Release candidate
Public Preview
v2.1.1 is similar to V2.1, and is available in all regions except: US, EU, and asia-southeast1.
For more information, see Managing processor versions.
Quotas and limits
Maximum pages (online/synchronous requests):
15
Maximum pages (batch/offline/asynchronous requests):
500
Maximum pages (imageless mode online/synchronous requests):
30
Note: Specific processor limits are in addition to
overall product quotas and limits .
Note: To extend the maximum page limit for online and synchronous
requests up to 30, be sure to enable imageless_mode in the
ProcessRequest .
Uptraining
Sample Input File
Open in new window.
Sample Output
Open in new window.
Supported regions
asia-south1
asia-southeast1
australia-southeast1
eu
europe-west2
europe-west3
northamerica-northeast1
us
Note: Consult Supported regions
for a full list of processor version availability by region.
More information
Enterprise Document OCR
Refer to Sample datasets
for sample labeled and unlabeled datasets to use for training.
Custom Extractor
Description
Extract fields from documents using generative AI or custom models; fine-tune models to accurately extract data from your documents.
Category
Extract
Functions
OCR, Entity Extraction
Release stage
General availability
Access status
Public
lock_open
Type in API
CUSTOM_EXTRACTION_PROCESSOR
Notes
Supported languages
Full list of languages
Language Name
BCP 47 Tag
Script
Handwriting supported
Afrikaans
af
Latn
Arabic
ar
Arab
Azerbaijani
az
Latn
Azerbaijani (Cyrillic)
az-Cyrl
Cyrl
Belarusian
be
Cyrl
Bulgarian
bg
Cyrl
Bosnian
bs
Latn
Catalan
ca
Latn
Cebuano
ceb
Latn
Czech
cs
Latn
Welsh
cy
Latn
Danish
da
Latn
German
de
Latn
Greek
el
Grek
English
en
Latn
Esperanto
eo
Latn
Spanish
es
Latn
Estonian
et
Latn
Basque
eu
Latn
Persian
fa
Arab
Finnish
fi
Latn
Filipino
fil
Latn
French
fr
Latn
Irish
ga
Latn
Galician
gl
Latn
Hindi
hi
Deva
Croatian
hr
Latn
Haitian Creole
ht
Latn
Hungarian
hu
Latn
Indonesian
id
Latn
Icelandic
is
Latn
Italian
it
Latn
Hebrew
iw
Hebr
Japanese
ja
Jpan
Javanese
jv
Latn
Kazakh
kk
Cyrl
Korean
ko
Kore
Kyrgyz
ky
Cyrl
Latin
la
Latn
Lithuanian
lt
Latn
Latvian
lv
Latn
Macedonian
mk
Cyrl
Mongolian
mn
Cyrl
Marathi
mr
Deva
Malay
ms
Latn
Maltese
mt
Latn
Nepali
ne
Deva
Dutch
nl
Latn
Norwegian
no
Latn
Polish
pl
Latn
Pashto
ps
Arab
Portuguese (Portugal & Brazil)
pt
Latn
Romanian
ro
Latn
Russian
ru
Cyrl
Russian (Petrine Orthography)
ru-PETR1708