OCR and document parsing models
Find OCR and document-parsing capabilities for images, scanned PDFs, tables, and complex layouts.
Map the topic before choosing
Separate text recognition, layout understanding, and full document parsing first. Then narrow candidates by Chinese support, table handling, local deployment, and throughput. Some OCR candidates are typed as models in the KG, so this page does not filter to tools only.
Use this page to understand the topic and discover candidates. When you are ready to compare deployment, budget, or capability constraints, open the decision guide above.
Curated KG entry points
A small set of entry points for discovering this topic, not a paid ranking or an official product review. Return to the homepage for the full semantic retrieval and second-stage summary.
PP-OCRv6
1.5M-parameter OCR model presented for low-latency edge deployment.
MechanismAgentic OCR
A multi-stage process running OCR, layout detection, and post-processing as sepa
ModelDeepSeek-OCR
🌟 Github 📥 Model Download 📄 Paper Link 📄 Arxiv Paper Link Explore the boundaries of visual-text compression.
ModelDeepSeek-OCR-2
🌟 Github 📥 Model Download 📄 Paper Link 📄 Arxiv Paper Link Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.9 + CUDA11.8:
ModelGLM-OCR
GLM-OCR is a multimodal OCR model for complex document understanding, built on the GLM-V encoder–decoder architecture. It introduces Multi-Token Prediction (MTP) loss and stable full-task reinforcement learning to improve training efficiency, recognition accuracy, and generalization. The model integrates the CogViT visual encoder pre-trained on large-scale image–text data, a lightweight cross-modal connector with efficient token downsampling, and a GLM-0.5B language decoder. Combined with a two-stage pipeline of la
ModelGOT-OCR-2.0-hf
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model - HF Transformers 🤗 implementation Haoran Wei*, Chenglong Liu*, Jinyue Chen, Jia Wang, Lingyu Kong, Yanming Xu, Zheng Ge, Liang Zhao, Jianjian Sun, Yuang Peng, Chunrui Han, Xiangyu Zhang
ModelLaTeX_OCR_rec
Hugging Face metadata identifies LaTeX_OCR_rec as image-to-text;PaddleOCR;image input;text output. No readable Model Card summary was captured; suitability still needs verification.
Modelnemotron-ocr-v1
The Nemotron OCR v1 model is a state-of-the-art text recognition model designed for robust end-to-end optical character recognition (OCR) on complex real-world images. It integrates three core neural network modules: a detector for text region localization, a recognizer for transcription of detected regions, and a relational model for layout and structure analysis.
Share feedback on this topic
Anonymous feedback is attached to this topic page, stays outside the KG, and is visible to other visitors.