Knowledge graph · semantic discovery

OCR 与文档解析模型

查找适合图片、扫描 PDF、表格和复杂版面的 OCR 与文档解析能力。

检索这个主题 →面向 AI 代理进入对应选型指南 →

公开 KG 候选 · 数据更新于 2026-09-09T19:04:28.116763+00:00 · 完整语义检索请回到首页使用

先看主题地图,再进入选型

先区分文字识别、版面理解和完整文档解析,再按中文支持、表格能力、本地部署与吞吐量缩小候选范围。当前 KG 中部分 OCR 候选被归为 model,因此主题页不会只筛选 tool。

先发现,再决策
本页用于理解主题范围、发现候选和相关概念;准备按部署、预算或能力约束比较时,请进入上方的对应选型指南。

主题精选 KG 入口

这里是用于发现本主题的一组精选入口,不代表付费排名,也不等同于官方产品评测;需要完整语义召回和二层总结,请回到首页检索。

模型

PP-OCRv6

1.5M-parameter OCR model presented for low-latency edge deployment.

机制

Agentic OCR

A multi-stage process running OCR, layout detection, and post-processing as sepa

模型

DeepSeek-OCR

来源摘要:🌟 Github 📥 Model Download 📄 Paper Link 📄 Arxiv Paper Link Explore the boundaries of visual-text compression.

模型

DeepSeek-OCR-2

来源摘要:🌟 Github 📥 Model Download 📄 Paper Link 📄 Arxiv Paper Link Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.9 + CUDA11.8:

模型

GLM-OCR

来源摘要:GLM-OCR is a multimodal OCR model for complex document understanding, built on the GLM-V encoder–decoder architecture. It introduces Multi-Token Prediction (MTP) loss and stable full-task reinforcement learning to improve training efficiency, recognition accuracy, and generalization. The model integrates the CogViT visual encoder pre-trained on large-scale image–text data, a lightweight cross-modal connector with efficient token downsampling, and a GLM-0.5B language decoder. Combined with a two-stage pipeline of la

模型

GOT-OCR-2.0-hf

来源摘要:General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model - HF Transformers 🤗 implementation Haoran Wei*, Chenglong Liu*, Jinyue Chen, Jia Wang, Lingyu Kong, Yanming Xu, Zheng Ge, Liang Zhao, Jianjian Sun, Yuang Peng, Chunrui Han, Xiangyu Zhang

模型

LaTeX_OCR_rec

Hugging Face 元数据记录:image-to-text;PaddleOCR;image input;text output。当前没有可读的模型卡片摘要,适用性仍需核验。

模型

nemotron-ocr-v1

来源摘要:The Nemotron OCR v1 model is a state-of-the-art text recognition model designed for robust end-to-end optical character recognition (OCR) on complex real-world images. It integrates three core neural network modules: a detector for text region localization, a recognizer for transcription of detected regions, and a relational model for layout and structure analysis.

留下这个主题的使用反馈

留言会关联到当前主题页面,保留在 KG 之外,其他访客也可以浏览和点赞。无需登录,不人工审核。

继续探索相关主题