awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

35 Repos

Awesome GitHub RepositoriesDocument Digitization Tools

Software for converting physical documents into searchable digital formats.

Distinguishing note: Focuses on the business process of digitization, distinct from general image processing.

Explore 35 awesome GitHub repositories matching business & productivity software · Document Digitization Tools. Refine with filters or upvote what's useful.

Awesome Document Digitization Tools GitHub Repositories

Finde die besten Repos mit KI.Wir suchen mit KI nach den am besten passenden Repositories.
  • awesome-selfhosted/awesome-selfhostedAvatar von awesome-selfhosted

    awesome-selfhosted/awesome-selfhosted

    299,516Auf GitHub ansehen↗

    Dieses Projekt ist ein von der Community kuratiertes Verzeichnis von Open-Source-Software, die für den Einsatz in privaten Serverumgebungen und Home-Labs konzipiert ist. Es dient als umfassende Ressource zur Entdeckung unabhängiger, selbst gehosteter Alternativen zu gängigen Cloud-Diensten und ermöglicht es Nutzern, die volle Datenhoheit und Kontrolle über ihre digitale Infrastruktur zu behalten. Das Verzeichnis ist durch eine hierarchische Taxonomie strukturiert, die eine riesige Sammlung von Anwendungen in logische Kategorien organisiert, von Medienmanagement und Datenanalyse bis hin zu privater Kommunikation und Tools für die Teamproduktivität. Es zeichnet sich durch einen kollaborativen Peer-Review-Prozess aus, bei dem Community-Mitglieder die Qualität und Relevanz jeder Einreichung validieren, um sicherzustellen, dass das Verzeichnis korrekt und zuverlässig bleibt. Das Projekt deckt ein breites Spektrum an Fähigkeiten ab, einschließlich Infrastruktur-Automatisierung, containerbasierter Service-Bereitstellung und deklarativem Konfigurationsmanagement. Diese Tools unterstützen Nutzer bei der Aufrechterhaltung reproduzierbarer Serverumgebungen und der Verwaltung komplexer Service-Abhängigkeiten auf privater Hardware. Das Verzeichnis wird als versionskontrolliertes Repository gepflegt, wodurch sichergestellt wird, dass alle Updates und Community-gesteuerten Änderungen nachverfolgt und transparent sind.

    Catalogs, converts, and serves electronic books to remote clients through a centralized interface.

    awesomeawesome-listcloud
    Auf GitHub ansehen↗299,516
  • ds4sd/doclingAvatar von DS4SD

    DS4SD/docling

    62,172Auf GitHub ansehen↗

    Docling is a multimodal content converter and document parser designed to transform PDFs, Office files, and HTML into structured Markdown or JSON for generative AI applications. It functions as an OCR document processor and a PDF layout analyzer that extracts tables, charts, and hierarchical structures while preserving the original page layout. The system operates as a local-first inference engine, allowing for the processing of sensitive data in air-gapped environments without external network connectivity. It can also be deployed as an API or a Model Context Protocol server to provide parsi

    Converts Office files, emails, and PDFs into unified formats like Markdown or JSON for organizational archival.

    Python
    Auf GitHub ansehen↗62,172
  • jaidedai/easyocrAvatar von JaidedAI

    JaidedAI/EasyOCR

    29,615Auf GitHub ansehen↗

    EasyOCR is a deep learning-based computer vision library designed to perform optical character recognition on images and video frames. It functions as a comprehensive pipeline that automates the transformation of visual text into machine-readable strings, enabling the digitization of physical documents, forms, and receipts into searchable data. The engine distinguishes itself through a multi-stage processing workflow that combines convolutional neural networks for spatial feature extraction with sequence-based decoding mechanisms. This architecture allows the system to identify and interpret

    Automates the conversion of physical paperwork into searchable digital text.

    Pythoncnncrnndata-mining
    Auf GitHub ansehen↗29,615
  • kovidgoyal/calibreAvatar von kovidgoyal

    kovidgoyal/calibre

    24,146Auf GitHub ansehen↗

    Calibre is a comprehensive suite for digital library management, serving as a centralized hub for organizing, converting, and editing e-book collections. It functions as a multi-purpose platform that combines a relational database for metadata tracking with a powerful processing engine capable of transforming document formats and restructuring internal markup. Beyond local management, the software acts as a content server, enabling users to host their libraries over a network for remote access and reading via standard web browsers. The project distinguishes itself through its deep extensibili

    Exposes local book collections over a network for remote browsing and reading via web browsers.

    Pythoncalibreebookebook-formats
    Auf GitHub ansehen↗24,146
  • datalab-to/suryaAvatar von datalab-to

    datalab-to/surya

    20,889Auf GitHub ansehen↗

    Surya is a document processing platform designed to transform unstructured files into structured, machine-readable data. It provides a comprehensive suite of tools for text recognition, layout analysis, and reading order detection, enabling the conversion of PDFs and images into formats such as JSON, HTML, or markdown. The platform is built to handle complex document workflows, offering capabilities for data extraction, document segmentation, and automated form completion. The platform distinguishes itself through a robust pipeline-based architecture that allows users to chain analysis tasks

    Programmatically injects structured data into PDF fields and visual overlays to streamline reporting.

    Python
    Auf GitHub ansehen↗20,889
  • wagtail/wagtailAvatar von wagtail

    wagtail/wagtail

    20,366Auf GitHub ansehen↗

    Wagtail is an open-source content management system built on the Django web framework. It provides a structured, tree-based approach to content modeling, allowing developers to define custom page types and reusable content components that are managed through a highly customizable administrative interface. The platform distinguishes itself through its flexible, block-based content composition system, which enables editors to assemble complex page layouts dynamically. It also offers robust support for multi-site and multi-lingual environments, allowing organizations to manage distinct websites

    Wagtail configures how files are delivered to users, balancing security permission checks against performance requirements through direct, redirected, or view-based serving methods.

    Pythoncmsdjangohacktoberfest
    Auf GitHub ansehen↗20,366
  • qwenlm/qwen2.5-vlAvatar von QwenLM

    QwenLM/Qwen2.5-VL

    19,480Auf GitHub ansehen↗

    Qwen2.5-VL ist ein autoregressiver multimodaler Transformer, der darauf ausgelegt ist, verschachtelte Sequenzen von Text- und visuellen Token zu verarbeiten. Er integriert visuelle Merkmalseinbettungen in einen gemeinsamen Sprachmodellraum, um modalübergreifendes Denken durchzuführen und kohärente Antworten oder strukturierten Layout-Code zu generieren. Das Projekt zeichnet sich durch Vision-Language-Action-Mapping aus, das es ihm ermöglicht, visuelle Schnittstellen wahrzunehmen und diese Wahrnehmung in umsetzbare Befehle für die Bedienung digitaler Bildschirme und Roboterhardware zu übersetzen. Es verwendet eine Bildkodierung mit dynamischer Auflösung und eine zeitliche Video-Indizierung, um diverse Bildgrößen und visuelle Sequenzen langer Dauer zu handhaben. Das Modell deckt ein breites Spektrum an Fähigkeiten ab, einschließlich mehrsprachiger optischer Zeichenerkennung für die Dokumentendigitalisierung, räumlicher Verankerung zur Lokalisierung von Objekten über Begrenzungsrahmen und der Analyse von Langform-Videoinhalten. Es unterstützt zudem multimodales mathematisches Denken, um Probleme mithilfe von Diagrammen und Grafiken zu lösen, und erweitert sein Verständnis auf eine Kontextlänge von einer Million Token.

    Converts images of scholarly papers and formulas into structured HTML or Markdown while preserving layout.

    Jupyter Notebook
    Auf GitHub ansehen↗19,480
  • scalar/scalarAvatar von scalar

    scalar/scalar

    13,979Auf GitHub ansehen↗

    Scalar is a platform for building and managing API specifications, focusing on OpenAPI and AsyncAPI standards. It provides tools to generate interactive API references with embedded testing interfaces, create mock servers for pre-implementation testing, and build offline-first API clients that sync with backend frameworks. The platform also supports version upgrades of specifications to maintain compatibility and includes command-line utilities for local development and document management. The project distinguishes itself through automated release workflows that generate changelogs and publi

    Serves OpenAPI documents locally via CLI with live file change detection and automatic refresh.

    TypeScriptapiapi-clientdocs
    Auf GitHub ansehen↗13,979
  • daybreak-u/chineseocr_liteAvatar von DayBreak-u

    DayBreak-u/chineseocr_lite

    12,324Auf GitHub ansehen↗

    chineseocr_lite is a lightweight Chinese optical character recognition engine designed to detect text regions, analyze orientation, and convert Chinese characters from images into digital text. It supports both horizontal and vertical reading layouts and can be deployed as a web service for image uploads and result visualization. The system utilizes a multi-backend inference framework that supports ncnn, mnn, and tnn, allowing it to run across diverse hardware and platforms. It is specifically engineered for lightweight deployment on mobile and desktop environments through the use of small mo

    Automates the extraction of structured text from images for integration into digital workflows.

    C++ncnnocrpytorch
    Auf GitHub ansehen↗12,324
  • cdnjs/cdnjsAvatar von cdnjs

    cdnjs/cdnjs

    10,707Auf GitHub ansehen↗

    cdnjs is a free, community-maintained content delivery network that hosts thousands of open-source frontend libraries. It delivers popular JavaScript and CSS assets from a global CDN to speed up website performance and reduce server load, with each library version stored as an immutable snapshot under a predictable directory structure. The platform provides a RESTful JSON API for programmatic access to library metadata, version details, and search functionality. This API returns structured data with HTTP cache headers, including immutable version details cached for nearly a year and library m

    Hosts and serves static assets for thousands of community-maintained open-source frontend libraries.

    cdncdnjscss
    Auf GitHub ansehen↗10,707
  • facebookresearch/nougatAvatar von facebookresearch

    facebookresearch/nougat

    10,015Auf GitHub ansehen↗

    Nougat is a neural OCR system and LLM document parser designed to convert images of academic PDF documents into structured markdown text and mathematical formulas. It functions as a PDF to markdown converter that uses deep learning to handle layout and formula recognition. The project provides a document training pipeline for generating datasets and training neural networks to recognize specific academic document styles. This includes utilities for training dataset generation, neural model training, and model checkpoint management to ensure reproducible deployment. The system covers a broad

    Digitizes scholarly papers into structured markdown while preserving mathematical formulas and complex tables.

    Python
    Auf GitHub ansehen↗10,015
  • akaunting/akauntingAvatar von akaunting

    akaunting/akaunting

    9,604Auf GitHub ansehen↗

    Akaunting is a modular business enterprise resource planning system and self-hosted accounting software. It provides a comprehensive platform for small business financial management, centering on a double-entry bookkeeping system with a general ledger and chart of accounts. The platform is designed for extensibility through a module-based architecture and a dedicated marketplace for procuring third-party applications. It supports multi-tenant data isolation and utilizes role-based access control to manage granular user permissions. Its capability surface covers a wide range of business opera

    Uploads images of receipts and expenses to capture financial data for conversion into records.

    PHPaccountingakauntingbalance
    Auf GitHub ansehen↗9,604
  • cvhub520/x-anylabelingAvatar von CVHub520

    CVHub520/X-AnyLabeling

    8,193Auf GitHub ansehen↗

    X-AnyLabeling is an AI-assisted annotation platform and computer vision labeling tool. It provides an interface for annotating images and videos using polygons and rectangles to create training sets for machine learning models. The project distinguishes itself through the integration of external AI models via a plugin-based inference backend, allowing for automated generation of candidate labels and the execution of specialized tasks like pose estimation and object detection. It also functions as an optical character recognition tool for extracting text and layout information from document im

    Converts document images into searchable digital formats by extracting text and layout information.

    Pythonartificial-intelligenceclipcomputer-vision
    Auf GitHub ansehen↗8,193
  • the-paperless-project/paperlessAvatar von the-paperless-project

    the-paperless-project/paperless

    7,917Auf GitHub ansehen↗

    Paperless is a self-hosted document management system designed to digitize, index, and archive paper documents. It functions as an optical character recognition system that converts scanned images and PDFs into a searchable digital library, providing a web-based interface for querying and retrieving documents from a database. The system features an automated file ingestion pipeline that monitors specific directories and email inboxes to process and import documents without manual uploading. To maintain a private archive, it includes on-disk encryption for sensitive files and the ability to or

    Converts physical papers into a searchable digital library through scanning and indexing for long-term storage.

    Python
    Auf GitHub ansehen↗7,917
  • rednote-hilab/dots.ocrAvatar von rednote-hilab

    rednote-hilab/dots.ocr

    7,695Auf GitHub ansehen↗

    dots.ocr is a suite of software utilities for document layout analysis, multilingual optical character recognition, and scene text digitization. It functions as an engine for extracting digital text and structured layout data from images and PDFs across various human scripts. The project includes a specialized transformer for converting charts, diagrams, and chemical formulas from raster images into scalable vector graphics. It also provides a pipeline to transform extracted text and structural layout from documents and web screenshots into formatted Markdown files. The system covers capabil

    Converts images and PDFs containing various human scripts into digital text while preserving document structure.

    Python
    Auf GitHub ansehen↗7,695
  • tesseract-ocr/tessdataAvatar von tesseract-ocr

    tesseract-ocr/tessdata

    7,586Auf GitHub ansehen↗

    This repository provides the pre-trained neural network and legacy data files used by Tesseract to recognize and extract printed text from images. It serves as a multilingual training data repository and a collection of Long Short-Term Memory models designed for high-accuracy optical character recognition across various global scripts and languages. The data includes specialized models for analyzing image layouts to determine text rotation and script direction. It provides the necessary language-specific datasets and linguistic patterns required to enable Tesseract OCR engines to function. T

    Provides the linguistic and visual data necessary for converting physical document scans into searchable digital formats.

    ocrtesseract
    Auf GitHub ansehen↗7,586
  • go-rod/rodAvatar von go-rod

    go-rod/rod

    6,713Auf GitHub ansehen↗

    Programmatically populates and submits web form fields for automation workflows.

    Goautomationcdpchrome-devtools
    Auf GitHub ansehen↗6,713
  • axa-group/parsrAvatar von axa-group

    axa-group/Parsr

    6,178Auf GitHub ansehen↗

    Parsr ist ein Extraktor für unstrukturierte Daten und eine Dokumenten-Parsing-Pipeline, die Rohdateien und Bilder in bereinigte, maschinenlesbare Formate konvertiert. Es fungiert als Dokumenten-Layout-Analysator und Pipeline zur Extraktion strukturierter Daten und Labels mittels Large Language Models. Das System enthält einen Dokumenten-Parsing-Visualizer, der ein grafisches Interface bietet, um Dokumente hochzuladen und den resultierenden strukturierten Datenausgang zu inspizieren. Das Projekt deckt Dokumentendigitalisierungs-Workflows ab, einschließlich Layout-Analyse zur Erkennung von Überschriften, Tabellen und Listen sowie automatisierte Dateneingabe durch die Bereinigung und Anreicherung unstrukturierter Inhalte.

    Converts physical or digital image files into clean text and organized data for downstream applications.

    JavaScript
    Auf GitHub ansehen↗6,178
  • lingui/js-linguiAvatar von lingui

    lingui/js-lingui

    5,786Auf GitHub ansehen↗

    Lingui is a JavaScript internationalization library that provides a framework-agnostic core with bindings for React, SolidJS, Svelte, Astro, and other JavaScript frameworks. It operates through a compile-time message extraction pipeline that scans source files for translatable strings, generates standard PO, JSON, or CSV catalog files, and compiles them into optimized JavaScript modules for production deployment. The library uses macro-based message definition to wrap translatable text in source code while preserving context for extraction, and includes a plural rule engine that automatically

    Serves documentation in a streamlined Markdown format optimized for AI consumption and auto-discovery.

    TypeScript
    Auf GitHub ansehen↗5,786
  • fontsource/fontsourceAvatar von fontsource

    fontsource/fontsource

    5,778Auf GitHub ansehen↗

    Serves documentation as plain Markdown files accessible by appending .md to any documentation URL.

    TypeScriptcssfontfont-family
    Auf GitHub ansehen↗5,778
Vorherige12Nächste
  1. Home
  2. Business & Productivity Software
  3. Document Digitization Tools

Unter-Tags erkunden

  • Accessibility DocumentationTechnical guides and design principles for building accessible digital products. **Distinct from Document Digitization Tools:** Distinct from general document digitization: focuses on accessibility-specific technical guidance rather than format conversion.
  • Document Serving4 Sub-TagsWeb daemons for hosting and delivering specialized document formats. **Distinct from Document Digitization Tools:** Distinct from Document Digitization: focuses on serving existing digital content rather than converting physical media.
  • Form AutomationTools for programmatically injecting data into digital form fields and overlays. **Distinct from Document Digitization Tools:** Distinct from general digitization: focuses on programmatic form population rather than just scanning.
  • Scholarly Document DigitizationSpecialized digitization of academic papers, focusing on formula and table preservation. **Distinct from Document Digitization Tools:** Focuses on scholarly content structures rather than general business document digitization.