نماذج OCR وتحليل المستندات
اكتشف قدرات التعرف على النصوص وتحليل المستندات للصور وملفات PDF الممسوحة والجداول والتخطيطات المعقدة.
ابدأ بالنطاق، ثم انتقل إلى القرار
ابدأ بالتمييز بين التعرف على النص، وفهم التخطيط، وتحليل المستند الكامل. ثم ضيّق المرشحين حسب دعم العربية والصينية، والجداول، والنشر المحلي، ومعدل المعالجة. بعض مرشحي OCR مصنفون كنماذج داخل الرسم المعرفي، لذلك لا تحصر هذه الصفحة النتائج في الأدوات.
هذه الصفحة تساعدك على فهم الموضوع واكتشاف المرشحين. عند الحاجة إلى مقارنة القيود والميزانية والنشر، اكتب المشكلة الفعلية في البحث الدلالي.
مداخل الرسم المعرفي
الأسماء والتفاصيل المعروضة هنا تبقى قريبة من السجل العام الأصلي عندما لا تتوفر ترجمة عربية موثقة؛ لا نملأ الفجوات بتخمينات.
PP-OCRv6
1.5M-parameter OCR model presented for low-latency edge deployment.
آليةAgentic OCR
A multi-stage process running OCR, layout detection, and post-processing as sepa
نموذجDeepSeek-OCR
🌟 Github 📥 Model Download 📄 Paper Link 📄 Arxiv Paper Link Explore the boundaries of visual-text compression.
نموذجDeepSeek-OCR-2
🌟 Github 📥 Model Download 📄 Paper Link 📄 Arxiv Paper Link Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.9 + CUDA11.8:
نموذجGLM-OCR
GLM-OCR is a multimodal OCR model for complex document understanding, built on the GLM-V encoder–decoder architecture. It introduces Multi-Token Prediction (MTP) loss and stable full-task reinforcement learning to improve training efficiency, recognition accuracy, and generalization. The model integrates the CogViT visual encoder pre-trained on large-scale image–text data, a lightweight cross-modal connector with efficient token downsampling, and a GLM-0.5B language decoder. Combined with a two-stage pipeline of la
نموذجGOT-OCR-2.0-hf
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model - HF Transformers 🤗 implementation Haoran Wei*, Chenglong Liu*, Jinyue Chen, Jia Wang, Lingyu Kong, Yanming Xu, Zheng Ge, Liang Zhao, Jianjian Sun, Yuang Peng, Chunrui Han, Xiangyu Zhang
نموذجLaTeX_OCR_rec
Hugging Face metadata identifies LaTeX_OCR_rec as image-to-text;PaddleOCR;image input;text output. No readable Model Card summary was captured; suitability still needs verification.
نموذجnemotron-ocr-v1
The Nemotron OCR v1 model is a state-of-the-art text recognition model designed for robust end-to-end optical character recognition (OCR) on complex real-world images. It integrates three core neural network modules: a detector for text region localization, a recognizer for transcription of detected regions, and a relational model for layout and structure analysis.
نموذجnemotron-ocr-v2
Nemotron OCR v2 is a state-of-the-art multilingual text recognition model designed for robust end-to-end optical character recognition (OCR) on complex real-world images. It integrates three core neural network modules: a detector for text region localization, a recognizer for transcription of detected regions, and a relational model for layout and structure analysis.
آليةOCR
Converts scanned PDFs into searchable text across 38 languages with up to 99% ac
نموذجQianfan-OCR
A Unified End-to-End Model for Document Intelligence **🤖 Demo** **📄 Technical Report** **🖥️ Qianfan Platform** **💻 GitHub** **🧩 Skill**
نموذجUnlimited-OCR
Transformers Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.3 + CUDA12.9: Please refer to the official vLLM recipe for deployment details
أضف تجربة إلى هذا الموضوع
الملاحظات العامة منفصلة عن حقائق الرسم المعرفي، ويمكن للزوار مناقشتها والتصويت عليها.
افتح مساحة النقاش ←