14 dépôts
Systems for preserving, managing, and cataloging digital collections.
Explore 14 awesome GitHub repositories matching part of an awesome list · Digital Archiving. Refine with filters or upvote what's useful.
ArchiveBox is a self-hosted archiving tool designed for personal digital preservation and research data management. It functions as an automated web preservation engine that monitors URL inputs from bookmarks, browser history, or manual entries to capture and store permanent, offline copies of web content. By utilizing headless browser automation, the system renders dynamic web pages to ensure that captured snapshots, PDFs, and media assets remain accurate and accessible even if the original source disappears. The project distinguishes itself through a modular extractor pipeline and a task-qu
Self-hosted wayback machine for archiving sites from various sources.
Dejavu is a Python audio fingerprinting library and recognition engine. It functions as a digital audio signature tool used to analyze sound waves and create unique identifiers for the purposes of audio search and retrieval. The project enables automatic music identification by matching live audio feeds or recorded clips against a database of fingerprints. It covers audio content matching and digital audio archiving to identify original source recordings from a stored collection. The system incorporates capabilities for generating audio fingerprints, identifying audio tracks, and recognizing
Supports digital audio archiving by generating fingerprints to organize and search large audio collections.
Lidarr est un gestionnaire d'automatisation de bibliothèque musicale et un contrôleur client qui surveille les artistes pour les nouvelles sorties et automatise l'acquisition de musique via BitTorrent et Usenet. Il sert d'organisateur de métadonnées et d'intégrateur, connectant les clients de téléchargement aux serveurs multimédias pour maintenir une collection de musique numérique complète et à jour. Le système se différencie par une maintenance automatisée de la bibliothèque, telle que la recherche de pistes manquantes pour combler les lacunes et la surveillance de versions de meilleure qualité des fichiers existants pour effectuer des mises à niveau automatiques de qualité. Il utilise des modèles de nommage configurables pour renommer les fichiers audio et organiser les dossiers, assurant une structure de système de fichiers local cohérente. Les capacités larges incluent la gestion automatisée des téléchargements, la recherche manuelle de sorties et la capacité de gérer les téléchargements échoués en cherchant des sorties alternatives. Le logiciel synchronise également les mises à jour de la bibliothèque et les changements de métadonnées avec les lecteurs multimédias et les serveurs.
Scans collections to find missing tracks and automatically upgrades files to higher quality bitrates.
CKAN is an open-source data management platform that provides the foundation for building data portals. It supports the full lifecycle of datasets—from creation and organization to publishing, cataloging with faceted search, and interactive data visualization—all through a web interface. The platform is built on a modular architecture that includes a plugin-based extensibility system, a harvesting framework for importing metadata from external sources, and a standardized RESTful JSON API for programmatic access to datasets and metadata. The web interface is rendered using the Jinja2 templatin
Tool for building and managing open data websites.
NAPS2 is a suite of document scanning software consisting of a desktop application, a command-line interface tool, and a networked scanner server. It serves as an interface for capturing images from scanners via TWAIN and WIA drivers, organizing those captures into digital documents, and exporting them to various file formats. The project distinguishes itself by providing a networked scanner server that shares local hardware across a network for remote image capture. It also includes a command-line tool for automating document capture and image processing workflows through scripts and termina
Captures physical pages, organizes their sequence, and saves them as digital files for long-term archiving.
An archiving tool with an IM-style interface that prioritizes privacy and accessibility, integrated with various archival services including Internet Archive, archive.today, Ghostarchive, IPFS, Telegraph, and file systems.
Toolkit for archiving webpages to various storage destinations.
Twitch VOD and Live Stream archiving platform. Includes a rendered and real-time chat for each archive.
Archiving platform for Twitch VODs and live streams.
Free and open-source digital preservation system designed to maintain standards-based, long-term access to collections of digital objects.
Digital preservation system for long-term access to collections.
Omeka S is a web publication system for universities, galleries, libraries, archives, and museums. It consists of a local network of independently curated exhibits sharing a collaboratively built pool of items, media, and their metadata.
Listed in the “Digital Archiving” section of the Awesome Selfhosted awesome list.
ArchivesSpace, the archives management tool
Management application for archives, manuscripts, and digital objects.
An automatic livestream recorder
Automatic recorder for capturing Twitch streams and metadata.
Cataloguing and data/media management application
Listed in the “Digital Archiving” section of the Awesome Selfhosted awesome list.
Open-source, web application for archival description and public access.
Standards-based archival description and access application.
Own webarchive service
Lightweight wayback machine for archiving bookmarks as HTML or PDF.