awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

21 个仓库

Awesome GitHub RepositoriesStructured Document Extraction

Processes that convert visual document layouts into machine-readable formats like JSON or Markdown.

Explore 21 awesome GitHub repositories matching artificial intelligence & ml · Structured Document Extraction. Refine with filters or upvote what's useful.

Awesome Structured Document Extraction GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • paddlepaddle/paddleocrPaddlePaddle 的头像

    PaddlePaddle/PaddleOCR

    82,412在 GitHub 上查看↗

    PaddleOCR is a comprehensive optical character recognition framework designed for detecting and transcribing text from images and documents into structured, machine-readable formats. It provides a modular computer vision pipeline that decouples image preprocessing, text detection, and character recognition into independent, configurable stages. This architecture supports automated document digitization and multilingual text recognition, capable of identifying text in over one hundred languages across diverse environments ranging from scanned documents to industrial scenes. The framework disti

    Transforms visual document layouts into structured, machine-readable formats like JSON or Markdown while correcting for perspective and artifacts.

    Pythonai4sciencechineseocrdocument-parsing
    在 GitHub 上查看↗82,412
  • opendatalab/mineruopendatalab 的头像

    opendatalab/MinerU

    67,734在 GitHub 上查看↗

    MinerU is a document parsing pipeline designed to transform unstructured files into machine-readable, structured data. It utilizes deep learning models to perform layout analysis, identifying document regions and extracting complex content such as mathematical expressions. By combining these neural network inferences with geometric heuristics, the system reconstructs the reading order and structural hierarchy of documents to ensure accurate data representation. The project distinguishes itself through a multi-stage processing workflow that integrates layout detection, optical character recogn

    Generates visual overlays that highlight detected text segments and reading order to verify parsing accuracy.

    Pythonai4sciencedocument-analysisextract-data
    在 GitHub 上查看↗67,734
  • vectifyai/pageindexVectifyAI 的头像

    VectifyAI/PageIndex

    33,103在 GitHub 上查看↗

    PageIndex is an agent-ready knowledge engine that processes documents into hierarchical tree structures to enable reasoning-based information retrieval. By organizing content into logical trees rather than relying on traditional vector database chunking, the platform preserves the original structure and flow of complex documents. It functions as a Model Context Protocol server, allowing external AI agents to connect to and query indexed knowledge bases through standardized communication protocols. The platform distinguishes itself by using vision-language models to process raw document images

    Processes raw document images directly to extract layout and structural information without relying on traditional OCR.

    Pythonagentagentic-aiai
    在 GitHub 上查看↗33,103
  • lightpanda-io/browserlightpanda-io 的头像

    lightpanda-io/browser

    31,168在 GitHub 上查看↗

    This project is a high-performance headless browser engine designed for scalable web automation, data extraction, and AI agent integration. It provides a specialized environment that allows autonomous agents and testing frameworks to interact with web content through standardized remote control protocols. By executing pages in a lightweight, headless state, the engine minimizes resource consumption while maintaining the ability to perform complex navigation and dynamic content rendering. The platform distinguishes itself through deep integration with AI-centric communication layers and advanc

    Parses the document structure into a simplified, machine-readable format to improve context for automated agents and language models.

    Zigbrowserbrowser-automationcdp
    在 GitHub 上查看↗31,168
  • opendataloader-project/opendataloader-pdfopendataloader-project 的头像

    opendataloader-project/opendataloader-pdf

    25,769在 GitHub 上查看↗

    This project is a PDF data extraction tool and document preprocessor designed to convert PDF files into structured formats such as Markdown, JSON, and HTML. It functions as an OCR document parser for scanned files, an accessibility automator for generating PDF/UA compliant metadata, and a loader for AI orchestration frameworks like LangChain. The software distinguishes itself through specialized handling of complex document elements, including the conversion of mathematical formulas into LaTeX and the generation of natural-language descriptions for charts and images. It utilizes recursive seg

    Overlays detected semantic elements onto original documents to visually verify and debug the extraction process.

    Javaa11yaccessibilityai
    在 GitHub 上查看↗25,769
  • cinnamon/kotaemonCinnamon 的头像

    Cinnamon/kotaemon

    25,139在 GitHub 上查看↗

    Kotaemon is an orchestration framework designed for building modular, agentic workflows that integrate document processing, retrieval-augmented generation, and multi-step reasoning. It provides a comprehensive platform for developing document-based question answering systems, allowing users to chain language models, prompt templates, and external tools into complex, automated pipelines. The system distinguishes itself through a highly modular architecture that emphasizes component-based composition and schema-driven data exchange. It supports autonomous agents capable of decomposing complex q

    Generates annotated debug images to verify the accuracy of document parsing and extraction.

    Pythonchatbotllmsopen-source
    在 GitHub 上查看↗25,139
  • microsoft/unilmmicrosoft 的头像

    microsoft/unilm

    22,030在 GitHub 上查看↗

    This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec

    Extracts information from structured documents like forms and receipts by analyzing both textual content and visual layout features.

    Pythonbeitbeit-3bitnet
    在 GitHub 上查看↗22,030
  • qwenlm/qwen2.5-vlQwenLM 的头像

    QwenLM/Qwen2.5-VL

    19,480在 GitHub 上查看↗

    Qwen2.5-VL 是一个自回归多模态 Transformer,旨在处理文本和视觉 Token 的交错序列。它将视觉特征嵌入集成到共享语言模型空间中,以执行跨模态推理并生成连贯的响应或结构化布局代码。 该项目通过视觉-语言-动作映射脱颖而出,使其能够感知视觉界面并将该感知转化为用于操作数字屏幕和机器人硬件的可执行命令。它采用动态分辨率图像编码和时间帧视频索引来处理不同的图像尺寸和长持续时间的视觉序列。 该模型涵盖了广泛的能力领域,包括用于文档数字化的多语言光学字符识别(OCR)、用于通过边界框定位对象的空间接地,以及长篇视频内容的分析。它还支持多模态数学推理以使用图表解决问题,并将理解能力扩展到一百万 Token 的上下文长度。

    Converts complex visual document layouts from research papers and magazines into structured machine-readable formats like Markdown.

    Jupyter Notebook
    在 GitHub 上查看↗19,480
  • qwenlm/qwen2-vlQwenLM 的头像

    QwenLM/Qwen2-VL

    19,404在 GitHub 上查看↗

    Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text, images, and video content. It functions as a visual reasoning engine and a visual agent framework, capable of interpreting visual data to perform object detection, document parsing, and spatial reasoning. The model is distinguished by its ability to act as a video understanding model, processing hour-long videos with second-level indexing and event recall. It further differentiates itself through a visual agent capability that interacts with software interfaces and robotic hardw

    Extracts text and structured data from documents and screenshots into machine-readable formats like HTML.

    Jupyter Notebook
    在 GitHub 上查看↗19,404
  • qwenlm/qwen3-vlQwenLM 的头像

    QwenLM/Qwen3-VL

    18,329在 GitHub 上查看↗

    Qwen3-VL is a multimodal vision-language model designed to process and reason across images, videos, and text. It functions as a computer vision framework capable of identifying objects, extracting structured data from documents, and interpreting spatial elements within visual media. The system operates as an automated user interface interaction agent, interpreting screen data to navigate software and mobile applications. By utilizing a unified transformer architecture, it performs complex visual reasoning to execute user-defined tasks without manual input. Beyond interface navigation, the m

    Converts multi-page documents into structured information using optical character recognition and contextual analysis.

    Jupyter Notebook
    在 GitHub 上查看↗18,329
  • unstructured-io/unstructuredUnstructured-IO 的头像

    Unstructured-IO/unstructured

    14,019在 GitHub 上查看↗

    Unstructured is an enterprise-grade data orchestration engine designed to transform raw, unstructured files into structured, machine-readable formats. It functions as a comprehensive platform for document ingestion, partitioning, and enrichment, specifically engineered to prepare complex data for retrieval-augmented generation and agentic AI workflows. The platform distinguishes itself through its sophisticated document processing strategies, which combine rule-based extraction with vision-language models to handle diverse file layouts, tables, and images. It provides a modular architecture t

    Uses vision-language models and rule-based strategies to parse complex document layouts into machine-readable JSON.

    HTMLdata-pipelinesdeep-learningdocument-image-analysis
    在 GitHub 上查看↗14,019
  • getomni-ai/zeroxgetomni-ai 的头像

    getomni-ai/zerox

    12,241在 GitHub 上查看↗

    Zerox is a multimodal document parser and OCR tool that uses vision models to convert PDF files and images into structured Markdown text. It functions as a visual layout extraction engine, leveraging large multimodal models to digitize documents while maintaining their original structural formatting. The system differentiates itself through the use of coordinate-based element mapping and multimodal layout analysis to identify structural elements like tables, charts, and headers. It utilizes rasterization to convert vector PDF pages into high-resolution bitmaps, ensuring consistent input for t

    Converts complex PDF files into structured Markdown while preserving tables, charts, and the original page formatting.

    TypeScriptocrpdf
    在 GitHub 上查看↗12,241
  • run-llama/liteparserun-llama 的头像

    run-llama/liteparse

    10,782在 GitHub 上查看↗

    A fast, helpful, and open-source document parser

    Converts PDFs and office documents into structured Markdown or JSON with spatial layout for direct use by language models.

    Rustdocument-ocrdocument-processingocr
    在 GitHub 上查看↗10,782
  • pymupdf/pymupdfpymupdf 的头像

    pymupdf/PyMuPDF

    9,086在 GitHub 上查看↗

    PyMuPDF is a comprehensive PDF manipulation library and document analysis tool. It serves as a text extraction tool, OCR engine, and image converter, providing a programmatic interface to edit, merge, split, and optimize PDF and Office documents. The project distinguishes itself through high-performance capabilities, including the use of C-bindings for low-level manipulation and parallelized page processing to accelerate workloads. It provides specialized conversion paths, such as transforming PDF content into Markdown for retrieval-augmented generation and large language model pipelines. It

    Converts visual document layouts into machine-readable formats like JSON, HTML, or XML.

    Pythondata-scienceepubextract-data
    在 GitHub 上查看↗9,086
  • bytedance/dolphinbytedance 的头像

    bytedance/Dolphin

    8,820在 GitHub 上查看↗

    Dolphin is a multimodal layout analyzer and image-to-structure converter that transforms photographed or digital document images into machine-readable structured data. It functions as an LLM document parser, utilizing vision-language models to simultaneously predict spatial layout and text content. The system is designed as a concurrent document processor, employing parallel document parsing to process multiple elements across distributed compute nodes. This high-throughput approach reduces the total time required to convert large volumes of images into structured formats. The project covers

    Converts visual document layouts into machine-readable formats like JSON or Markdown.

    Pythondocument-analysislayout-analysisocr
    在 GitHub 上查看↗8,820
  • kreuzberg-dev/kreuzbergkreuzberg-dev 的头像

    kreuzberg-dev/kreuzberg

    8,527在 GitHub 上查看↗

    Kreuzberg is a document extraction engine that converts PDFs, Office files, images, and over 90 other formats into clean, structured text and metadata. It is built around a compiled Rust core that can be used as a native library, a command-line tool, a REST API server, or a WebAssembly module for browser-based processing. The system is designed to run entirely on self-hosted infrastructure, with no data leaving the user's environment. What distinguishes Kreuzberg is its breadth of integration surfaces and its pipeline architecture. It exposes extraction capabilities through native bindings fo

    Returns a traversable tree of nodes with heading levels and inline annotations for knowledge graphs.

    Rustdocument-intelligenceelixirffi
    在 GitHub 上查看↗8,527
  • mindee/doctrmindee 的头像

    mindee/doctr

    6,149在 GitHub 上查看↗

    DocTR is a deep learning OCR library built on PyTorch that detects and transcribes text in document images using a two-stage detection-recognition pipeline. It provides a complete framework for building and deploying OCR pipelines with pretrained models available through the Hugging Face Hub, and supports exporting trained models to ONNX format for cross-runtime deployment. The library offers end-to-end OCR pipelines that combine text detection and recognition to extract all text from document images or PDFs, with support for rotated page handling and varied text orientations. It includes cap

    Structures detected text into a hierarchy of words, lines, blocks, pages, and documents.

    Pythondeep-learningdocument-recognitionocr
    在 GitHub 上查看↗6,149
  • layout-parser/layout-parserLayout-Parser 的头像

    Layout-Parser/layout-parser

    5,749在 GitHub 上查看↗

    Layout-parser 是一个深度学习文档布局解析器和图像分析框架。它提供了一个工具包,用于从扫描文档和数字图像中提取结构信息和布局模式,将其转换为程序化数据结构以进行自动化分析。 该框架将布局检测与光学字符识别(OCR)集成,将表格区域转换为机器可读数据。它利用神经网络来识别和分类文档图像中的结构元素,而不依赖于手动基于规则的系统。 该系统涵盖了广泛的文档分析功能,包括文档结构解析、自动化表格提取和分层布局表示。它还包括可视化工具,用于在原始图像上渲染检测到的元素和层次结构,以进行结果验证。

    Organizes detected document elements into a parent-child tree structure to preserve logical information flow.

    Python
    在 GitHub 上查看↗5,749
  • katanaml/sparrowkatanaml 的头像

    katanaml/sparrow

    5,162在 GitHub 上查看↗

    Sparrow 是一个 LLM 文档提取平台和基于视觉的推理引擎,旨在将图像和 PDF 转换为经过验证的结构化数据。它作为代理工作流编排器,将分类、提取和验证任务串联成多步流水线。 该系统的特色在于其后端无关的推理层,可管理本地 GPU、Apple Silicon 和云服务商上的模型。它利用基于坐标的视觉定位将提取的文本映射到精确的边界框坐标,并使用基于提示的模型引导来引导注意力并规范化数据格式。 该平台涵盖了文档智能工作流,包括用于保持结构完整性的专业图像表格处理,以及用于验证提取字段正确性的模式驱动验证。它还提供了一个用于监控 API 性能、使用分析和系统健康状况的文档分析仪表板。 架构包含一个基于插件的扩展系统,用于集成索引和编排中使用的第三方库。

    Uses vision-capable language models to parse document layouts and convert visual content into structured data.

    Pythonagentic-aicomputer-visiondocumentai
    在 GitHub 上查看↗5,162
  • datalab-to/chandradatalab-to 的头像

    datalab-to/chandra

    4,833在 GitHub 上查看↗

    sChandra is a document processing platform that converts images, PDFs, Word documents, spreadsheets, and other formats into structured output such as HTML, Markdown, or JSON while preserving layout. It can also extract specific data fields from invoices, contracts, or reports using user-defined JSON schemas, with citations back to source locations. The service supports form filling in PDF and image documents, document generation from Markdown, and extraction of tracked changes from Word files. The platform distinguishes itself with pipeline-based processing chains that combine multiple proces

    Converts PDFs, images, and Office files into structured HTML, Markdown, or JSON while preserving layout.

    Pythonaiocr
    在 GitHub 上查看↗4,833
上一个12下一个
  1. Home
  2. Artificial Intelligence & ML
  3. Natural Language Processing
  4. Structured Document Extraction

探索子标签

  • Heading Level ClassifiersCapabilities for identifying heading levels (H1-H6) from font size clustering and semantic analysis in documents. **Distinct from Structured Document Extraction:** Distinct from Structured Document Extraction: focuses specifically on classifying heading hierarchy rather than general layout-to-markdown conversion.
  • Hierarchical RepresentationsReturns a traversable tree of nodes with parent-child references, heading levels, and inline annotations for knowledge graphs. **Distinct from Structured Document Extraction:** Distinct from Structured Document Extraction: focuses on the hierarchical tree representation with heading levels and annotations, not just conversion to JSON or Markdown.
  • Tree RepresentationsRepresents a document as a flat array of nodes with index-based parent/child references forming a tree. **Distinct from Structured Document Extraction:** Distinct from Structured Document Extraction: focuses on the internal tree representation of the document structure rather than the conversion of visual layouts to machine-readable formats.
  • Vision-Language Model BackendsUses vision-language models as an OCR backend and for extracting structured JSON from documents using a schema. **Distinct from Structured Document Extraction:** Distinct from Structured Document Extraction: focuses on using VLM backends for extraction, not general layout-to-text conversion.
  • Visual Debugging UtilitiesTools that generate visual overlays to verify the accuracy of automated document parsing and text detection.