Knowledge graph · semantic discovery

OCR and document parsing models

Find OCR and document-parsing capabilities for images, scanned PDFs, tables, and complex layouts.

Search this topic →For AI agentsOpen decision guide →

Public KG candidates · updated 2026-09-09T19:04:28.116763+00:00 · semantic retrieval remains available on the main site

Map the topic before choosing

Separate text recognition, layout understanding, and full document parsing first. Then narrow candidates by Chinese support, table handling, local deployment, and throughput. Some OCR candidates are typed as models in the KG, so this page does not filter to tools only.

Discovery first
Use this page to understand the topic and discover candidates. When you are ready to compare deployment, budget, or capability constraints, open the decision guide above.

Curated KG entry points

A small set of entry points for discovering this topic, not a paid ranking or an official product review. Return to the homepage for the full semantic retrieval and second-stage summary.

Model

PP-OCRv6

1.5M-parameter OCR model presented for low-latency edge deployment.

Mechanism

Agentic OCR

A multi-stage process running OCR, layout detection, and post-processing as sepa

Model

DeepSeek-OCR

🌟 Github 📥 Model Download 📄 Paper Link 📄 Arxiv Paper Link Explore the boundaries of visual-text compression.

Model

DeepSeek-OCR-2

🌟 Github 📥 Model Download 📄 Paper Link 📄 Arxiv Paper Link Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.9 + CUDA11.8:

Model

GLM-OCR

GLM-OCR is a multimodal OCR model for complex document understanding, built on the GLM-V encoder–decoder architecture. It introduces Multi-Token Prediction (MTP) loss and stable full-task reinforcement learning to improve training efficiency, recognition accuracy, and generalization. The model integrates the CogViT visual encoder pre-trained on large-scale image–text data, a lightweight cross-modal connector with efficient token downsampling, and a GLM-0.5B language decoder. Combined with a two-stage pipeline of la

Model

GOT-OCR-2.0-hf

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model - HF Transformers 🤗 implementation Haoran Wei*, Chenglong Liu*, Jinyue Chen, Jia Wang, Lingyu Kong, Yanming Xu, Zheng Ge, Liang Zhao, Jianjian Sun, Yuang Peng, Chunrui Han, Xiangyu Zhang

Model

LaTeX_OCR_rec

Hugging Face metadata identifies LaTeX_OCR_rec as image-to-text;PaddleOCR;image input;text output. No readable Model Card summary was captured; suitability still needs verification.

Model

nemotron-ocr-v1

The Nemotron OCR v1 model is a state-of-the-art text recognition model designed for robust end-to-end optical character recognition (OCR) on complex real-world images. It integrates three core neural network modules: a detector for text region localization, a recognizer for transcription of detected regions, and a relational model for layout and structure analysis.

Share feedback on this topic

Anonymous feedback is attached to this topic page, stays outside the KG, and is visible to other visitors.

Explore related topics