GLM-OCR is a multimodal OCR model for complex document understanding, built on the GLM-V encoder–decoder architecture. It introduces Multi-Token Prediction (MTP) loss and stable full-task reinforcement learning to improve training efficiency, recognition accuracy, and generalization. The model integrates the CogViT visual encoder pre-trained on large-scale image–text data, a lightweight cross-modal connector with efficient token downsampling, and a GLM-0.5B language decoder. Combined with a two-stage pipeline of layout analysis and parallel recognition based on PP-DocLayout-V3, GLM-OCR delivers robust and high-quality OCR performance across diverse document layouts.
Key Features
State-of-the-Art Performance
: Achieves a score of 94.62 on OmniDocBench V1.5, ranking #1 overall, and delivers state-of-the-art results across major document understanding benchmarks, including formula recognition, table recognition, and information extraction.
Optimized for Real-World Scenarios
: Designed and optimized for practical business use cases, maintaining robust performance on complex tables, code-heavy documents, seals, and other challenging real-world layouts.
Efficient Inference
: With only 0.9B parameters, GLM-OCR supports deployment via vLLM, SGLang, and Ollama, significantly reducing inference latency and compute cost, making it ideal for high-concurrency services and edge deployments.
Easy to Use
: Fully open-sourced and equipped with a comprehensive
SDK
and inference toolchain, offering simple installation, one-line invocation, and smooth integration into existing production pipelines.
Information Extraction
– extract structured information from documents. Prompts must follow a strict JSON schema. For example, to extract personal ID information:
The GLM-OCR model is released under the MIT License.
The complete OCR pipeline integrates
PP-DocLayoutV3
for document layout analysis, which is licensed under the Apache License 2.0. Users should comply with both licenses when using this project.
Runs of zai-org GLM-OCR on huggingface.co
1.9M
Total runs
0
24-hour runs
14.6K
3-day runs
-108.7K
7-day runs
-1.3M
30-day runs
More Information About GLM-OCR huggingface.co Model
GLM-OCR huggingface.co is an AI model on huggingface.co that provides GLM-OCR's model effect (), which can be used instantly with this zai-org GLM-OCR model. huggingface.co supports a free trial of the GLM-OCR model, and also provides paid use of the GLM-OCR. Support call GLM-OCR model through api, including Node.js, Python, http.
GLM-OCR huggingface.co is an online trial and call api platform, which integrates GLM-OCR's modeling effects, including api services, and provides a free online trial of GLM-OCR, you can try GLM-OCR online for free by clicking the link below.
zai-org GLM-OCR online free url in huggingface.co:
GLM-OCR is an open source model from GitHub that offers a free installation service, and any user can find GLM-OCR on GitHub to install. At the same time, huggingface.co provides the effect of GLM-OCR install, users can directly use GLM-OCR installed effect in huggingface.co for debugging and trial. It also supports api for free installation.