GLM-OCR
A multimodal OCR model for reading complex documents and turning images of text into usable text. It targets accurate and fast document understanding.
Share on XLicense: Apache-2.0
Overview
GLM-OCR is a multimodal OCR model for understanding complex documents. It combines a visual encoder, a small cross-modal connector and a GLM-0.5B language decoder, and uses a two-stage pipeline of layout analysis followed by parallel recognition. It turns document images into text, including formulas and tables, and can be deployed with vLLM, SGLang or Ollama.
Key features
- Two-stage layout analysis and parallel recognition
- Handles formulas, tables and information extraction
- Only 0.9B parameters
- Deployable with vLLM, SGLang and Ollama
- Aimed at complex tables, code-heavy documents and seals
Best for
Teams that need to read complex real-world documents into text and want a compact model they can host themselves.
- Upstream
- zai-org/GLM-OCR
- Fork on GitHub
- Guo-astro/GLM-OCR
- Upstream stars
- 7.5k
- Category
- Documents, media and content
- Language
- Python
- License
- Apache-2.0
- Forked
- 2026-03-17
- Sync status
- In syncLast synced 2026-10-10
More in Documents, media and content
Generate spoken audio from text on a CPU with Pocket TTS, a lightweight Python text-to-speech tool. Use its Python API or CLI without a GPU or speech-generation web API.
Forked 2026-10-09Last synced 2026-10-10License: MITDocuments, media and contentGitHub
A local FFmpeg toolkit for editing and checking video and audio with coding agents or directly, without uploading media or using API keys.
Forked 2026-10-04Last synced 2026-10-10License: MITDocuments, media and contentGitHub
Chinese translation of a deep learning textbook, making the book easier to read and improve through community feedback.
Forked 2026-10-04Last synced 2026-10-10No license declaredDocuments, media and contentGitHub
A Markdown-based typesetting system that turns plain text into papers, presentations, websites, books and knowledge bases. It adds scripting features to Markdown for richer documents.
Forked 2026-09-28Last synced 2026-10-10License: GPL-3.0Documents, media and contentGitHub