For ai driven document processing, the strongest matches are deepseek-ai/deepseek-ocr (This framework provides the core OCR and visual encoding), katanaml/sparrow (Sparrow is a self-hostable, LLM-driven platform that orchestrates document) and vikparuchuri/marker (Marker is an AI-powered document processing tool that uses). unstructured-io/unstructured and microsoft/markitdown round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
Explore the best AI-driven document processing tools. Compare top open-source libraries by features and activity to find the best fit for your project.
DeepSeek-OCR is a vision processing framework designed to convert image-based text into machine-readable tokens for large language models. It functions as a document inference pipeline that encodes visual data into compact representations, enabling automated optical character recognition and document analysis workflows. The system distinguishes itself through a high-throughput architecture that utilizes hardware-accelerated batch inference to process large volumes of visual data. It incorporates dynamic resolution scaling to manage the balance between visual detail and token consumption, ensu
This framework provides the core OCR and visual encoding pipeline necessary for AI-driven document extraction, though it functions as a specialized component for feeding multimodal models rather than a complete, end-to-end document processing application.
Sparrow is an LLM document extraction platform and vision-based inference engine designed to convert images and PDFs into validated structured data. It functions as an agentic workflow orchestrator that chains classification, extraction, and validation tasks into multi-step pipelines. The system distinguishes itself through a backend-agnostic inference layer that manages models across local GPUs, Apple Silicon, and cloud providers. It employs coordinate-based visual grounding to map extracted text to precise bounding box coordinates and utilizes hint-based model steering to guide attention an
Sparrow is a self-hostable, LLM-driven platform that orchestrates document classification and structured data extraction from images and PDFs, directly matching the requirements for an AI-powered document processing system.
Marker is an LLM-powered document parser and OCR pipeline designed to convert PDFs and unstructured files into structured markdown, JSON, and HTML. It functions as a data preprocessor that transforms complex documents into machine-readable formats while preserving tables, equations, and layout structures. The system utilizes large language models to refine OCR accuracy, clean mathematical notation, and merge fragmented tables across multiple pages. It employs model-based layout analysis to predict block types and bounding boxes, ensuring a more precise conversion of document elements. Capabi
Marker is an AI-powered document processing tool that uses LLMs and OCR to convert unstructured PDFs into structured formats, fitting the core requirements for data extraction and document parsing.
Unstructured is an enterprise-grade data orchestration engine designed to transform raw, unstructured files into structured, machine-readable formats. It functions as a comprehensive platform for document ingestion, partitioning, and enrichment, specifically engineered to prepare complex data for retrieval-augmented generation and agentic AI workflows. The platform distinguishes itself through its sophisticated document processing strategies, which combine rule-based extraction with vision-language models to handle diverse file layouts, tables, and images. It provides a modular architecture t
Unstructured is a comprehensive platform designed to ingest, partition, and extract structured data from diverse document formats using a combination of OCR and vision-language models, making it a direct fit for AI-powered document processing.
This project is an AI-powered document processing engine designed to transform diverse file formats into structured Markdown. By leveraging multimodal language models, it performs complex layout analysis and semantic text extraction, allowing for the conversion of both unstructured files and scanned images into machine-readable content. The toolkit distinguishes itself through a modular, plugin-based architecture that orchestrates multi-stage extraction pipelines. Users can steer the parsing behavior by injecting custom instructions, enabling the system to adapt to domain-specific document st
This tool provides an AI-driven pipeline for converting diverse, unstructured document formats into structured Markdown, effectively serving as a foundational engine for data extraction and document intelligence tasks.
docetl is an AI-powered document ETL tool and map-reduce orchestrator designed to transform large collections of unstructured documents into structured, queryable tables using language models. It provides a declarative pipeline framework for extracting, cleaning, and transforming data from sources such as PDFs and text files into predefined schemas. The project distinguishes itself through a semantic data integration suite that enables joining datasets and resolving duplicate entities based on embedding-based similarity. It includes an interactive prompt playground for developing and optimizi
docetl is a self-hostable, AI-powered ETL framework that uses LLMs to extract, classify, and structure data from unstructured documents into queryable schemas, directly matching the requirements for automated document processing.
Marker is a comprehensive document processing platform designed to automate the conversion, extraction, and structuring of data from complex files. It functions as an orchestration engine that chains modular processing steps into versioned, reusable pipelines, allowing organizations to standardize document handling and automate repetitive business tasks at scale. The platform distinguishes itself through its support for secure, private infrastructure deployment, enabling users to run containerized services within their own environments to maintain strict data privacy. It features specialized
Marker is a self-hostable document processing platform that provides the necessary orchestration and extraction engines to convert unstructured files into structured data, fitting the core requirements for automated document handling.
llmware is a Python framework for AI agent orchestration and model management, designed to coordinate multi-model workflows and autonomous agents. It provides a unified model catalog and standardized interface to execute specialized language models for complex research, analysis, and structured data generation. The project distinguishes itself through its heavy emphasis on local execution and quantized inference, allowing models to run on private infrastructure using CPU, GPU, and NPU acceleration via runtimes like ONNX and OpenVino. It features a specialized ability to translate natural lang
This framework provides the necessary tools for document extraction, classification, and structured data generation using local LLMs, making it a capable foundation for building AI-powered document processing pipelines.
Zerox is a multimodal document parser and OCR tool that uses vision models to convert PDF files and images into structured Markdown text. It functions as a visual layout extraction engine, leveraging large multimodal models to digitize documents while maintaining their original structural formatting. The system differentiates itself through the use of coordinate-based element mapping and multimodal layout analysis to identify structural elements like tables, charts, and headers. It utilizes rasterization to convert vector PDF pages into high-resolution bitmaps, ensuring consistent input for t
Zerox is a multimodal document parser that uses vision models to extract and structure data from PDFs into Markdown, fitting the core requirement for AI-powered document processing even though it focuses on layout-aware conversion rather than specific field-level data extraction.
Surya is a document processing platform designed to transform unstructured files into structured, machine-readable data. It provides a comprehensive suite of tools for text recognition, layout analysis, and reading order detection, enabling the conversion of PDFs and images into formats such as JSON, HTML, or markdown. The platform is built to handle complex document workflows, offering capabilities for data extraction, document segmentation, and automated form completion. The platform distinguishes itself through a robust pipeline-based architecture that allows users to chain analysis tasks
Surya provides a robust pipeline for document segmentation, layout analysis, and text extraction that serves as a core engine for AI-powered data processing, though it focuses more on the structural analysis and OCR components than on high-level LLM-based semantic parsing.
Kreuzberg is a document extraction engine that converts PDFs, Office files, images, and over 90 other formats into clean, structured text and metadata. It is built around a compiled Rust core that can be used as a native library, a command-line tool, a REST API server, or a WebAssembly module for browser-based processing. The system is designed to run entirely on self-hosted infrastructure, with no data leaving the user's environment. What distinguishes Kreuzberg is its breadth of integration surfaces and its pipeline architecture. It exposes extraction capabilities through native bindings fo
Kreuzberg is a self-hostable document extraction engine that provides the necessary OCR, multi-format support, and API-first architecture to structure unstructured files, though it focuses more on extraction and metadata than on LLM-based classification.
Kotaemon is an orchestration framework designed for building modular, agentic workflows that integrate document processing, retrieval-augmented generation, and multi-step reasoning. It provides a comprehensive platform for developing document-based question answering systems, allowing users to chain language models, prompt templates, and external tools into complex, automated pipelines. The system distinguishes itself through a highly modular architecture that emphasizes component-based composition and schema-driven data exchange. It supports autonomous agents capable of decomposing complex q
Kotaemon is an orchestration framework for building document-based RAG and reasoning pipelines that includes robust tools for data extraction, table parsing, and OCR integration, making it a strong foundation for custom document processing workflows.
A fast, helpful, and open-source document parser
This tool provides a comprehensive suite for document parsing, OCR, and data extraction that supports API-first integration and self-hosted deployment, making it a direct fit for AI-powered document processing pipelines.
MinerU is a document parsing pipeline designed to transform unstructured files into machine-readable, structured data. It utilizes deep learning models to perform layout analysis, identifying document regions and extracting complex content such as mathematical expressions. By combining these neural network inferences with geometric heuristics, the system reconstructs the reading order and structural hierarchy of documents to ensure accurate data representation. The project distinguishes itself through a multi-stage processing workflow that integrates layout detection, optical character recogn
MinerU is a document parsing pipeline that uses deep learning for layout analysis and OCR to convert unstructured documents into structured formats, fitting the core requirements for AI-powered data extraction.
Docling is a modular framework designed for document parsing, layout analysis, and structured data extraction. It transforms unstructured files and web content into a unified, hierarchical data model that preserves the spatial and semantic relationships between text, tables, images, and layout elements. By normalizing diverse input formats into a consistent internal representation, the library enables uniform processing across various document types. The project distinguishes itself through a schema-driven approach that maps document regions to strongly-typed objects, ensuring data accuracy t
Docling is a powerful framework for parsing and structuring complex document layouts into machine-readable formats, providing the essential extraction and normalization capabilities needed for AI-driven document processing pipelines.
Langextract is a framework designed to transform unstructured text into structured, machine-readable data using language model orchestration. It provides a high-performance pipeline that processes large volumes of narrative text by utilizing parallel execution and sequential extraction passes. The library is built to handle complex data extraction tasks, including specialized support for clinical information and medical entity relationship recognition. The project distinguishes itself through a plugin-based architecture that supports both local hardware execution and cloud-hosted model endpoi
This framework provides the necessary orchestration and pipeline capabilities to extract and structure data from unstructured text using language models, though it functions as a developer-focused library rather than a standalone document processing application.
This platform serves as a comprehensive environment for managing private language models, document knowledge bases, and automated agent workflows within secure local infrastructure. It functions as a document-aware workspace that enables users to ingest diverse file formats into searchable repositories, ensuring that all data processing and model inference remain within private, local environments to maintain data sovereignty. The system distinguishes itself through a modular agentic engine that allows for the definition of custom skills and external tool execution. By utilizing a multi-model
This platform provides a document-aware workspace that handles ingestion, parsing, and vector-based retrieval for unstructured files, making it a capable tool for document-based AI workflows even though it focuses more on RAG and chat than on structured data extraction.
PaddleOCR is a comprehensive optical character recognition framework designed for detecting and transcribing text from images and documents into structured, machine-readable formats. It provides a modular computer vision pipeline that decouples image preprocessing, text detection, and character recognition into independent, configurable stages. This architecture supports automated document digitization and multilingual text recognition, capable of identifying text in over one hundred languages across diverse environments ranging from scanned documents to industrial scenes. The framework disti
PaddleOCR is a powerful OCR and document structure analysis framework that provides the essential text extraction and layout parsing capabilities required for document processing, though it relies on modular integration rather than being an all-in-one LLM-based extraction platform.
This project is a document digitization utility that combines traditional optical character recognition with language model processing to convert scanned PDF files into structured markdown. It functions as an automated pipeline that extracts raw text from images and applies intelligent post-processing to refine the output. The system distinguishes itself by using language models to perform error correction, removing artifacts and formatting inconsistencies common in raw character recognition. It incorporates a modular design that decouples processing logic from specific model providers, allow
This tool uses LLMs to refine and structure OCR output from scanned documents, providing a specialized pipeline for extracting and formatting data that aligns with the core requirements of AI-powered document processing.
Parsr is an unstructured data extractor and document parsing pipeline that converts raw files and images into cleaned, machine-readable formats. It functions as a document layout analyzer and a pipeline for extracting structured data and labels using large language models. The system includes a document parsing visualizer, providing a graphical interface to upload documents and inspect the resulting structured data output. The project covers document digitization workflows, including layout analysis to detect headings, tables, and lists, and automated data entry through the cleaning and enri
Parsr is a document parsing pipeline that handles OCR, layout analysis, and structured data extraction, providing the core functionality needed for AI-powered document processing even if it focuses more on layout-based extraction than pure LLM-based parsing.
Grobid is a machine learning system designed to transform academic and scientific PDF publications into structured XML. It functions as a PDF to XML parser and scholarly metadata extractor, identifying and normalizing titles, authors, affiliations, and bibliographic references from research papers. The system utilizes a deep learning document segmenter to divide raw PDFs into functional regions and employs a bibliographic reference resolver to match citations against external registries for metadata enrichment and DOI resolution. It supports a full machine learning model training pipeline, al
Grobid is a specialized machine learning system for extracting structured data from scientific PDFs, providing the core document parsing and metadata extraction capabilities required for AI-powered document processing.
PageIndex is an agent-ready knowledge engine that processes documents into hierarchical tree structures to enable reasoning-based information retrieval. By organizing content into logical trees rather than relying on traditional vector database chunking, the platform preserves the original structure and flow of complex documents. It functions as a Model Context Protocol server, allowing external AI agents to connect to and query indexed knowledge bases through standardized communication protocols. The platform distinguishes itself by using vision-language models to process raw document images
PageIndex is a document processing and indexing engine that uses vision-language models to structure unstructured content into hierarchical trees for AI agent retrieval, fitting the core requirements for AI-powered data extraction and document organization.
| Repository | Stele | Limbaj | Licență | Ultimul push |
|---|---|---|---|---|
| deepseek-ai/deepseek-ocr | 22.5K | Python | mit | |
| katanaml/sparrow | 5.2K | Python | GPL-3.0 | |
| vikparuchuri/marker | 36.2K | Python | GPL-3.0 | |
| unstructured-io/unstructured | 14K | HTML | apache-2.0 | |
| microsoft/markitdown | 154.5K | Python | MIT | |
| ucbepic/docetl | 3.6K | Python | mit | |
| datalab-to/marker | 36.1K | Python | GPL-3.0 | |
| llmware-ai/llmware | 14.8K | Python | Apache-2.0 | |
| getomni-ai/zerox | 12.2K | TypeScript | MIT | |
| datalab-to/surya | 20.9K | Python | Apache-2.0 |