21 مستودعات
Processes that convert visual document layouts into machine-readable formats like JSON or Markdown.
Explore 21 awesome GitHub repositories matching artificial intelligence & ml · Structured Document Extraction. Refine with filters or upvote what's useful.
PaddleOCR is a comprehensive optical character recognition framework designed for detecting and transcribing text from images and documents into structured, machine-readable formats. It provides a modular computer vision pipeline that decouples image preprocessing, text detection, and character recognition into independent, configurable stages. This architecture supports automated document digitization and multilingual text recognition, capable of identifying text in over one hundred languages across diverse environments ranging from scanned documents to industrial scenes. The framework disti
Transforms visual document layouts into structured, machine-readable formats like JSON or Markdown while correcting for perspective and artifacts.
MinerU is a document parsing pipeline designed to transform unstructured files into machine-readable, structured data. It utilizes deep learning models to perform layout analysis, identifying document regions and extracting complex content such as mathematical expressions. By combining these neural network inferences with geometric heuristics, the system reconstructs the reading order and structural hierarchy of documents to ensure accurate data representation. The project distinguishes itself through a multi-stage processing workflow that integrates layout detection, optical character recogn
Generates visual overlays that highlight detected text segments and reading order to verify parsing accuracy.
PageIndex is an agent-ready knowledge engine that processes documents into hierarchical tree structures to enable reasoning-based information retrieval. By organizing content into logical trees rather than relying on traditional vector database chunking, the platform preserves the original structure and flow of complex documents. It functions as a Model Context Protocol server, allowing external AI agents to connect to and query indexed knowledge bases through standardized communication protocols. The platform distinguishes itself by using vision-language models to process raw document images
Processes raw document images directly to extract layout and structural information without relying on traditional OCR.
This project is a high-performance headless browser engine designed for scalable web automation, data extraction, and AI agent integration. It provides a specialized environment that allows autonomous agents and testing frameworks to interact with web content through standardized remote control protocols. By executing pages in a lightweight, headless state, the engine minimizes resource consumption while maintaining the ability to perform complex navigation and dynamic content rendering. The platform distinguishes itself through deep integration with AI-centric communication layers and advanc
Parses the document structure into a simplified, machine-readable format to improve context for automated agents and language models.
This project is a PDF data extraction tool and document preprocessor designed to convert PDF files into structured formats such as Markdown, JSON, and HTML. It functions as an OCR document parser for scanned files, an accessibility automator for generating PDF/UA compliant metadata, and a loader for AI orchestration frameworks like LangChain. The software distinguishes itself through specialized handling of complex document elements, including the conversion of mathematical formulas into LaTeX and the generation of natural-language descriptions for charts and images. It utilizes recursive seg
Overlays detected semantic elements onto original documents to visually verify and debug the extraction process.
Kotaemon is an orchestration framework designed for building modular, agentic workflows that integrate document processing, retrieval-augmented generation, and multi-step reasoning. It provides a comprehensive platform for developing document-based question answering systems, allowing users to chain language models, prompt templates, and external tools into complex, automated pipelines. The system distinguishes itself through a highly modular architecture that emphasizes component-based composition and schema-driven data exchange. It supports autonomous agents capable of decomposing complex q
Generates annotated debug images to verify the accuracy of document parsing and extraction.
This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec
Extracts information from structured documents like forms and receipts by analyzing both textual content and visual layout features.
Qwen2.5-VL هو محول متعدد الوسائط ذاتي الانحدار مصمم لمعالجة تسلسلات متداخلة من الرموز النصية والمرئية. يدمج تضمينات الميزات المرئية في مساحة نموذج لغوي مشترك لإجراء استدلال متعدد الوسائط وتوليد استجابات متماسكة أو كود تخطيط مهيكل. يتميز المشروع برسم خرائط الرؤية-اللغة-العمل، مما يسمح له بإدراك الواجهات المرئية وترجمة هذا الإدراك إلى أوامر قابلة للتنفيذ لتشغيل الشاشات الرقمية وأجهزة الروبوت. يستخدم ترميز الصور بدقة ديناميكية وفهرسة الفيديو ذات الإطارات الزمنية للتعامل مع أحجام الصور المتنوعة وتسلسلات الفيديو طويلة المدة. يغطي النموذج نطاقاً واسعاً من القدرات، بما في ذلك التعرف الضوئي على الحروف متعدد اللغات لرقمنة المستندات، والتأريض المكاني لتحديد موقع الكائنات عبر مربعات الإحاطة، وتحليل محتوى الفيديو طويل الشكل. كما يدعم الاستدلال الرياضي متعدد الوسائط لحل المشكلات باستخدام المخططات والرسوم البيانية، ويمتد فهمه إلى طول سياق يبلغ مليون رمز.
Converts complex visual document layouts from research papers and magazines into structured machine-readable formats like Markdown.
Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text, images, and video content. It functions as a visual reasoning engine and a visual agent framework, capable of interpreting visual data to perform object detection, document parsing, and spatial reasoning. The model is distinguished by its ability to act as a video understanding model, processing hour-long videos with second-level indexing and event recall. It further differentiates itself through a visual agent capability that interacts with software interfaces and robotic hardw
Extracts text and structured data from documents and screenshots into machine-readable formats like HTML.
Qwen3-VL is a multimodal vision-language model designed to process and reason across images, videos, and text. It functions as a computer vision framework capable of identifying objects, extracting structured data from documents, and interpreting spatial elements within visual media. The system operates as an automated user interface interaction agent, interpreting screen data to navigate software and mobile applications. By utilizing a unified transformer architecture, it performs complex visual reasoning to execute user-defined tasks without manual input. Beyond interface navigation, the m
Converts multi-page documents into structured information using optical character recognition and contextual analysis.
Unstructured is an enterprise-grade data orchestration engine designed to transform raw, unstructured files into structured, machine-readable formats. It functions as a comprehensive platform for document ingestion, partitioning, and enrichment, specifically engineered to prepare complex data for retrieval-augmented generation and agentic AI workflows. The platform distinguishes itself through its sophisticated document processing strategies, which combine rule-based extraction with vision-language models to handle diverse file layouts, tables, and images. It provides a modular architecture t
Uses vision-language models and rule-based strategies to parse complex document layouts into machine-readable JSON.
Zerox is a multimodal document parser and OCR tool that uses vision models to convert PDF files and images into structured Markdown text. It functions as a visual layout extraction engine, leveraging large multimodal models to digitize documents while maintaining their original structural formatting. The system differentiates itself through the use of coordinate-based element mapping and multimodal layout analysis to identify structural elements like tables, charts, and headers. It utilizes rasterization to convert vector PDF pages into high-resolution bitmaps, ensuring consistent input for t
Converts complex PDF files into structured Markdown while preserving tables, charts, and the original page formatting.
A fast, helpful, and open-source document parser
Converts PDFs and office documents into structured Markdown or JSON with spatial layout for direct use by language models.
PyMuPDF is a comprehensive PDF manipulation library and document analysis tool. It serves as a text extraction tool, OCR engine, and image converter, providing a programmatic interface to edit, merge, split, and optimize PDF and Office documents. The project distinguishes itself through high-performance capabilities, including the use of C-bindings for low-level manipulation and parallelized page processing to accelerate workloads. It provides specialized conversion paths, such as transforming PDF content into Markdown for retrieval-augmented generation and large language model pipelines. It
Converts visual document layouts into machine-readable formats like JSON, HTML, or XML.
Dolphin is a multimodal layout analyzer and image-to-structure converter that transforms photographed or digital document images into machine-readable structured data. It functions as an LLM document parser, utilizing vision-language models to simultaneously predict spatial layout and text content. The system is designed as a concurrent document processor, employing parallel document parsing to process multiple elements across distributed compute nodes. This high-throughput approach reduces the total time required to convert large volumes of images into structured formats. The project covers
Converts visual document layouts into machine-readable formats like JSON or Markdown.
Kreuzberg is a document extraction engine that converts PDFs, Office files, images, and over 90 other formats into clean, structured text and metadata. It is built around a compiled Rust core that can be used as a native library, a command-line tool, a REST API server, or a WebAssembly module for browser-based processing. The system is designed to run entirely on self-hosted infrastructure, with no data leaving the user's environment. What distinguishes Kreuzberg is its breadth of integration surfaces and its pipeline architecture. It exposes extraction capabilities through native bindings fo
Returns a traversable tree of nodes with heading levels and inline annotations for knowledge graphs.
DocTR is a deep learning OCR library built on PyTorch that detects and transcribes text in document images using a two-stage detection-recognition pipeline. It provides a complete framework for building and deploying OCR pipelines with pretrained models available through the Hugging Face Hub, and supports exporting trained models to ONNX format for cross-runtime deployment. The library offers end-to-end OCR pipelines that combine text detection and recognition to extract all text from document images or PDFs, with support for rotated page handling and varied text orientations. It includes cap
Structures detected text into a hierarchy of words, lines, blocks, pages, and documents.
Layout-parser هو إطار عمل للتعلم العميق لتحليل تخطيط المستندات وتحليل الصور. يوفر مجموعة أدوات لاستخراج المعلومات الهيكلية وأنماط التخطيط من المستندات الممسوحة ضوئياً والصور الرقمية، وتحويلها إلى هياكل بيانات برمجية للتحليل الآلي. يدمج إطار العمل اكتشاف التخطيط مع التعرف الضوئي على الحروف لتحويل المناطق الجدولية إلى بيانات مقروءة آلياً. يستخدم الشبكات العصبية لتحديد وتصنيف العناصر الهيكلية داخل صور المستندات دون الاعتماد على أنظمة يدوية قائمة على القواعد. يغطي النظام مجموعة واسعة من قدرات تحليل المستندات، بما في ذلك تحليل هيكل المستند، واستخراج الجدول الآلي، وتمثيل التخطيط الهرمي. يتضمن أيضاً أدوات تصور لرسم العناصر المكتشفة والتسلسلات الهرمية فوق الصور الأصلية للتحقق من النتائج.
Organizes detected document elements into a parent-child tree structure to preserve logical information flow.
Sparrow هي منصة لاستخراج البيانات من المستندات تعتمد على النماذج اللغوية الكبيرة (LLM) ومحرك استنتاج بصري مصمم لتحويل الصور وملفات PDF إلى بيانات مهيكلة وموثقة. تعمل كمنسق لسير عمل الوكلاء (agentic workflow) الذي يربط مهام التصنيف والاستخراج والتحقق في خطوط معالجة متعددة الخطوات. يتميز النظام بطبقة استنتاج مستقلة عن الخلفية (backend-agnostic) تدير النماذج عبر وحدات معالجة الرسوميات المحلية، وApple Silicon، ومزودي الخدمات السحابية. يستخدم النظام التحديد البصري القائم على الإحداثيات لربط النص المستخرج بإحداثيات دقيقة، ويستخدم توجيه النماذج القائم على التلميحات لتوجيه الانتباه وتوحيد تنسيقات البيانات. تغطي المنصة سير عمل ذكاء المستندات، بما في ذلك معالجة الجداول القائمة على الصور للحفاظ على السلامة الهيكلية، والتحقق القائم على المخططات لضمان صحة الحقول المستخرجة. كما توفر لوحة تحكم لتحليل المستندات لمراقبة أداء واجهة برمجة التطبيقات (API) وتحليلات الاستخدام وصحة النظام. تتضمن البنية نظام إضافات (plugin-based) لدمج مكتبات الطرف الثالث المستخدمة في الفهرسة والتنسيق.
Uses vision-capable language models to parse document layouts and convert visual content into structured data.
sChandra is a document processing platform that converts images, PDFs, Word documents, spreadsheets, and other formats into structured output such as HTML, Markdown, or JSON while preserving layout. It can also extract specific data fields from invoices, contracts, or reports using user-defined JSON schemas, with citations back to source locations. The service supports form filling in PDF and image documents, document generation from Markdown, and extraction of tracked changes from Word files. The platform distinguishes itself with pipeline-based processing chains that combine multiple proces
Converts PDFs, images, and Office files into structured HTML, Markdown, or JSON while preserving layout.