awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

8 Repos

Awesome GitHub RepositoriesDigital Preservation Tools

Utilities for capturing and maintaining long-term accessible versions of web content.

Distinguishing note: Specifically targets the creation of multiple preservation-grade formats like PDFs and screenshots for long-term access.

Explore 8 awesome GitHub repositories matching content management & publishing · Digital Preservation Tools. Refine with filters or upvote what's useful.

Awesome Digital Preservation Tools GitHub Repositories

Finde die besten Repos mit KI.Wir suchen mit KI nach den am besten passenden Repositories.
  • awesome-selfhosted/awesome-selfhostedAvatar von awesome-selfhosted

    awesome-selfhosted/awesome-selfhosted

    299,516Auf GitHub ansehen↗

    Dieses Projekt ist ein von der Community kuratiertes Verzeichnis von Open-Source-Software, die für den Einsatz in privaten Serverumgebungen und Home-Labs konzipiert ist. Es dient als umfassende Ressource zur Entdeckung unabhängiger, selbst gehosteter Alternativen zu gängigen Cloud-Diensten und ermöglicht es Nutzern, die volle Datenhoheit und Kontrolle über ihre digitale Infrastruktur zu behalten. Das Verzeichnis ist durch eine hierarchische Taxonomie strukturiert, die eine riesige Sammlung von Anwendungen in logische Kategorien organisiert, von Medienmanagement und Datenanalyse bis hin zu privater Kommunikation und Tools für die Teamproduktivität. Es zeichnet sich durch einen kollaborativen Peer-Review-Prozess aus, bei dem Community-Mitglieder die Qualität und Relevanz jeder Einreichung validieren, um sicherzustellen, dass das Verzeichnis korrekt und zuverlässig bleibt. Das Projekt deckt ein breites Spektrum an Fähigkeiten ab, einschließlich Infrastruktur-Automatisierung, containerbasierter Service-Bereitstellung und deklarativem Konfigurationsmanagement. Diese Tools unterstützen Nutzer bei der Aufrechterhaltung reproduzierbarer Serverumgebungen und der Verwaltung komplexer Service-Abhängigkeiten auf privater Hardware. Das Verzeichnis wird als versionskontrolliertes Repository gepflegt, wodurch sichergestellt wird, dass alle Updates und Community-gesteuerten Änderungen nachverfolgt und transparent sind.

    Maintains modular storage systems for the long-term archiving and secure dissemination of digital library assets.

    awesomeawesome-listcloud
    Auf GitHub ansehen↗299,516
  • hiroi-sora/umi-ocrAvatar von hiroi-sora

    hiroi-sora/Umi-OCR

    45,273Auf GitHub ansehen↗

    Umi-OCR is an optical character recognition engine designed to convert visual text from images and documents into machine-readable character data. It functions as a local-first toolkit, processing all visual data directly on the host machine using embedded neural network models to maintain privacy and offline availability. The project distinguishes itself through its focus on automated document digitization and integrated barcode and QR code decoding. By utilizing a modular, Python-based orchestration layer, it enables users to transform static image files and multi-page documents into search

    Converts large volumes of scanned documents or images into searchable text files automatically.

    Pythonocrocr-pythonpaddleocr
    Auf GitHub ansehen↗45,273
  • paperless-ngx/paperless-ngxAvatar von paperless-ngx

    paperless-ngx/paperless-ngx

    42,172Auf GitHub ansehen↗

    Paperless-ngx is a self-hosted document management server designed to transform physical paperwork into a searchable, organized digital archive. It functions as a private platform for storing, indexing, and retrieving documents, providing users with full control over their data on local infrastructure or private cloud servers. The system distinguishes itself through an automated workflow engine that categorizes, tags, and routes incoming files using content analysis and metadata extraction. To maintain responsiveness during resource-intensive tasks like optical character recognition, it utili

    Builds a searchable digital repository for physical paperwork through automation.

    Pythonangulararchivingdjango
    Auf GitHub ansehen↗42,172
  • xinntao/real-esrganAvatar von xinntao

    xinntao/Real-ESRGAN

    35,798Auf GitHub ansehen↗

    Real-ESRGAN is a deep learning restoration pipeline designed to enhance low-resolution media and improve the visual quality of damaged photographs. It functions as a generative image upscaler that reconstructs high-resolution details from source inputs by utilizing neural networks trained to fill in missing information and remove noise. The project distinguishes itself as a blind super-resolution tool, meaning it improves image sharpness and fidelity without requiring prior knowledge of the specific degradation applied to the source. It employs high-order degradation modeling to address compl

    Digitizes and enhances historical or low-quality media assets to ensure they remain clear for modern viewing standards.

    Pythonaminedenoiseesrgan
    Auf GitHub ansehen↗35,798
  • pirate/archiveboxAvatar von pirate

    pirate/ArchiveBox

    27,721Auf GitHub ansehen↗

    ArchiveBox is a self-hosted web archiving system designed to capture and preserve permanent static copies of webpages, media, and PDFs on personal infrastructure. It functions as a digital content curator and personal web archive manager, allowing users to import URLs from bookmarks, RSS feeds, and browser history to create a centralized, searchable knowledge base. The project is distinguished by its ability to archive private, paywalled, or login-protected content using browser cookies and authenticated session persistence. It ensures long-term availability by saving pages in multiple concur

    Captures and maintains long-term accessible versions of web content in multiple preservation-grade formats like PDF and PNG.

    Python
    Auf GitHub ansehen↗27,721
  • squidfunk/mkdocs-materialAvatar von squidfunk

    squidfunk/mkdocs-material

    26,949Auf GitHub ansehen↗

    This project is a comprehensive documentation site framework and static site generator theme designed to transform markdown files into professional, responsive websites. It functions as a technical content platform that supports complex documentation projects, including multi-project management, blog workflows, and advanced content formatting. By processing source files through an extensible pipeline, it generates self-contained HTML sites that can be hosted on any web server without a database. What distinguishes this framework is its focus on developer experience and highly configurable bui

    The documentation generator configures archive pages to display collections of posts by date, including options for pagination, custom naming, and URL formatting.

    Pythondocumentationframeworkmaterial-design
    Auf GitHub ansehen↗26,949
  • archivebox/archiveboxAvatar von ArchiveBox

    ArchiveBox/ArchiveBox

    26,876Auf GitHub ansehen↗

    ArchiveBox is a self-hosted archiving tool designed for personal digital preservation and research data management. It functions as an automated web preservation engine that monitors URL inputs from bookmarks, browser history, or manual entries to capture and store permanent, offline copies of web content. By utilizing headless browser automation, the system renders dynamic web pages to ensure that captured snapshots, PDFs, and media assets remain accurate and accessible even if the original source disappears. The project distinguishes itself through a modular extractor pipeline and a task-qu

    Create multiple versions of web pages including screenshots, PDFs, and media files to ensure that content remains readable and accessible for long-term digital preservation and reference.

    Pythonarchiveboxbackupsbookmark-archiver
    Auf GitHub ansehen↗26,876
  • ruffle-rs/ruffleAvatar von ruffle-rs

    ruffle-rs/ruffle

    18,187Auf GitHub ansehen↗

    Ruffle is an Adobe Flash Player emulator built to execute legacy animation and interactive content within modern web browsers and desktop environments. By utilizing a high-performance WebAssembly engine, it interprets legacy bytecode and scripting languages to render content without requiring original plugins or outdated software. The project functions as a cross-platform media player that preserves access to archived digital assets by simulating the original runtime environment. The emulator distinguishes itself through its ability to automatically detect and replace obsolete media objects o

    Identifies and automatically replaces obsolete media objects on websites to restore functionality and ensure long-term accessibility.

    Rustemulatorflashhacktoberfest
    Auf GitHub ansehen↗18,187
  1. Home
  2. Content Management & Publishing
  3. Content Archiving
  4. Digital Preservation Tools

Unter-Tags erkunden

  • Web Asset Preservation ToolsUtilities for restoring and preserving access to obsolete web media objects. **Distinct from Digital Preservation Tools:** Focuses on restoring functionality to obsolete web assets rather than general digital archiving.