awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

13 dépôts

Awesome GitHub RepositoriesData Management Systems

Platforms for document storage, search, and business intelligence.

Explore 13 awesome GitHub repositories matching part of an awesome list · Data Management Systems. Refine with filters or upvote what's useful.

Awesome Data Management Systems GitHub Repositories

Trouvez les meilleurs dépôts grâce à l'IA.Nous recherchons les dépôts les plus pertinents grâce à l'IA.
  • apache/incubator-supersetAvatar de apache

    apache/incubator-superset

    73,325Voir sur GitHub↗

    This project is a business intelligence suite and SQL data visualization platform used for data analysis, reporting, and monitoring. It provides a web application for exploring datasets and building interactive dashboards, complemented by a web-based SQL query editor for analyzing raw data from connected stores. The platform features a semantic data layer to define standardized metrics and dimensions, ensuring consistent data interpretation across reports. It includes a security framework with role-based access control to manage user permissions and authentication across shared dashboards. T

    Web application for data exploration and visualization.

    TypeScript
    Voir sur GitHub↗73,325
  • ocrmypdf/ocrmypdfAvatar de ocrmypdf

    ocrmypdf/OCRmyPDF

    33,898Voir sur GitHub↗

    OCRmyPDF is a command-line tool designed to transform scanned documents into searchable, selectable PDF files. It functions as a document processing pipeline that adds a hidden text layer to image-based files while simultaneously optimizing the document's file size and image quality. By preserving the original visual fidelity of the input, it ensures that digitized documents remain accessible to screen readers and search engines. The project distinguishes itself through a modular architecture that supports custom plugins and the integration of external recognition engines, allowing users to t

    Adds searchable text layers to scanned PDF documents.

    Pythonimage-processingocrpdf
    Voir sur GitHub↗33,898
  • getredash/redashAvatar de getredash

    getredash/redash

    28,653Voir sur GitHub↗

    Redash is a self-hosted analytics platform and SQL data visualization tool. It provides a web-based SQL query editor for writing, executing, and scheduling database queries, and functions as a business intelligence dashboard for monitoring metrics via visual widgets. The platform distinguishes itself through its data source connectors, which integrate with various SQL, NoSQL, and API-based stores to retrieve information for analysis. It enables self-service analytics by allowing users to run queries with dynamic parameters and supports shared data reporting via public links or embedded dashbo

    Dashboard construction and data visualization for business intelligence.

    Pythonanalyticsathenabi
    Voir sur GitHub↗28,653
  • pirate/archiveboxAvatar de pirate

    pirate/ArchiveBox

    27,721Voir sur GitHub↗

    ArchiveBox is a self-hosted web archiving system designed to capture and preserve permanent static copies of webpages, media, and PDFs on personal infrastructure. It functions as a digital content curator and personal web archive manager, allowing users to import URLs from bookmarks, RSS feeds, and browser history to create a centralized, searchable knowledge base. The project is distinguished by its ability to archive private, paywalled, or login-protected content using browser cookies and authenticated session persistence. It ensures long-term availability by saving pages in multiple concur

    Self-hosted web archive for creating browsable local backups.

    Python
    Voir sur GitHub↗27,721
  • iterative/dvcAvatar de iterative

    iterative/dvc

    15,680Voir sur GitHub↗

    DVC is a data versioning tool and pipeline orchestrator designed to track large datasets and machine learning models. It functions as a system for managing large data artifacts by storing lightweight metadata in version control while keeping the actual binaries in a separate cache. The project serves as an experiment tracker and remote storage synchronizer, enabling the execution and comparison of machine learning iterations based on hyperparameters and performance metrics. It provides a bridge for pushing and pulling these large data artifacts between local environments and cloud or on-premi

    Version control system for data in machine learning projects.

    Python
    Voir sur GitHub↗15,680
  • the-paperless-project/paperlessAvatar de the-paperless-project

    the-paperless-project/paperless

    7,917Voir sur GitHub↗

    Paperless is a self-hosted document management system designed to digitize, index, and archive paper documents. It functions as an optical character recognition system that converts scanned images and PDFs into a searchable digital library, providing a web-based interface for querying and retrieving documents from a database. The system features an automated file ingestion pipeline that monitors specific directories and email inboxes to process and import documents without manual uploading. To maintain a private archive, it includes on-disk encryption for sensitive files and the ability to or

    System for indexing and archiving scanned paper documents.

    Python
    Voir sur GitHub↗7,917
  • beancount/beancountAvatar de beancount

    beancount/beancount

    5,291Voir sur GitHub↗

    Beancount is a plain-text double-entry accounting system. It enforces zero-sum transactions, organizes accounts into a hierarchical five-type tree, and verifies balances at specific dates using precision-derived tolerances. Transactions are recorded in plain-text files with a strict syntax that supports currency-specific rounding, automatic interpolation of missing amounts, and comprehensive metadata including tags, links, and payee annotations. Beyond core bookkeeping, Beancount offers investment portfolio tracking with lot-based cost basis management, configurable booking strategies (FIFO,

    Double-entry bookkeeping language for financial records.

    Pythonbeancount
    Voir sur GitHub↗5,291
  • bram2w/baserowAvatar de bram2w

    bram2w/baserow

    5,085Voir sur GitHub↗

    Baserow est une base de données relationnelle no-code et un constructeur d'applications qui permet aux utilisateurs de créer des tables de données structurées et des outils métier via une interface visuelle. Il fonctionne comme un backend de données API REST headless et un espace de travail de données auto-hébergé, fournissant une plateforme pour gérer des bases de données collaboratives tout en conservant un contrôle total sur la résidence des données. La plateforme intègre des modèles de langage étendus pour servir de plateforme de données alimentée par LLM, capable de générer des structures de base de données, du contenu d'enregistrement et des flux de travail techniques à partir du langage naturel. Elle agit également comme un serveur de protocole de contexte de modèle (Model Context Protocol), permettant aux agents IA distants d'interagir avec des enregistrements de base de données structurés par programmation. Au-delà de ses capacités fondamentales de base de données, le projet fournit des outils pour construire des portails externes de marque, des applications métier internes et des tableaux de bord interactifs. Il inclut un moteur d'automatisation piloté par les événements pour l'automatisation des processus métier et prend en charge une large gamme d'intégrations API, incluant les webhooks, le streaming d'événements WebSocket et la synchronisation de données tierces. Le logiciel est conçu pour l'hébergement sur infrastructure privée et le déploiement conteneurisé afin d'assurer la souveraineté et la sécurité des données.

    No-code persistence platform combining databases and spreadsheets.

    Python
    Voir sur GitHub↗5,085
  • mathesar-foundation/mathesarAvatar de mathesar-foundation

    mathesar-foundation/mathesar

    5,012Voir sur GitHub↗

    Mathesar is a no-code database manager and PostgreSQL GUI that provides a visual interface for managing relational database structures and records. It functions as a low-code data platform for administering schemas, tables, and relationships without the need to write manual SQL commands. The platform allows for the creation of shareable forms to collect data and the management of file attachments linked directly to database records. It includes a PostgreSQL administration tool for controlling database roles, user permissions, and data validation rules. The system covers relational data model

    Spreadsheet-like interface for PostgreSQL databases.

    Svelteairtable-alternativeautomatic-apidatabase-access
    Voir sur GitHub↗5,012
  • beancount/favaAvatar de beancount

    beancount/fava

    2,499Voir sur GitHub↗

    Fava is a web-based dashboard and query tool for visualizing and analyzing financial records stored in Beancount plain-text ledger files. It serves as a double-entry bookkeeping viewer and plain-text accounting dashboard that renders ledger files as interactive reports, searchable financial tables, and visual tools for exploring balance sheets and income statements. The project distinguishes itself through a specialized BQL query interface that executes SQL-like queries against postings to extract specific financial data and trends. It includes a financial data visualization system for genera

    Web interface for managing double-entry bookkeeping records.

    Pythonbeancountledgerplaintext-accounting
    Voir sur GitHub↗2,499
  • rd17/ambarAvatar de RD17

    RD17/ambar

    1,949Voir sur GitHub↗

    :mag: Ambar: Document Search Engine

    Document search engine with automated crawling and OCR.

    JavaScript
    Voir sur GitHub↗1,949
  • camelot-dev/excaliburAvatar de camelot-dev

    camelot-dev/excalibur

    1,790Voir sur GitHub↗

    Web interface for extracting tabular data from PDF files.

    Pythonextractfor-humanspdf
    Voir sur GitHub↗1,790
  • artefactual/archivematicaAvatar de artefactual

    artefactual/archivematica

    506Voir sur GitHub↗

    Free and open-source digital preservation system designed to maintain standards-based, long-term access to collections of digital objects.

    Digital preservation system for long-term access to collections.

    Pythonarchivematicadigital-preservation
    Voir sur GitHub↗506
  1. Home
  2. Part of an Awesome List
  3. Databases & Data
  4. Data Management Systems