awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

13 个仓库

Awesome GitHub RepositoriesData Management Systems

Platforms for document storage, search, and business intelligence.

Explore 13 awesome GitHub repositories matching part of an awesome list · Data Management Systems. Refine with filters or upvote what's useful.

Awesome Data Management Systems GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • apache/incubator-supersetapache 的头像

    apache/incubator-superset

    73,325在 GitHub 上查看↗

    This project is a business intelligence suite and SQL data visualization platform used for data analysis, reporting, and monitoring. It provides a web application for exploring datasets and building interactive dashboards, complemented by a web-based SQL query editor for analyzing raw data from connected stores. The platform features a semantic data layer to define standardized metrics and dimensions, ensuring consistent data interpretation across reports. It includes a security framework with role-based access control to manage user permissions and authentication across shared dashboards. T

    Web application for data exploration and visualization.

    TypeScript
    在 GitHub 上查看↗73,325
  • ocrmypdf/ocrmypdfocrmypdf 的头像

    ocrmypdf/OCRmyPDF

    33,898在 GitHub 上查看↗

    OCRmyPDF is a command-line tool designed to transform scanned documents into searchable, selectable PDF files. It functions as a document processing pipeline that adds a hidden text layer to image-based files while simultaneously optimizing the document's file size and image quality. By preserving the original visual fidelity of the input, it ensures that digitized documents remain accessible to screen readers and search engines. The project distinguishes itself through a modular architecture that supports custom plugins and the integration of external recognition engines, allowing users to t

    Adds searchable text layers to scanned PDF documents.

    Pythonimage-processingocrpdf
    在 GitHub 上查看↗33,898
  • getredash/redashgetredash 的头像

    getredash/redash

    28,653在 GitHub 上查看↗

    Redash is a self-hosted analytics platform and SQL data visualization tool. It provides a web-based SQL query editor for writing, executing, and scheduling database queries, and functions as a business intelligence dashboard for monitoring metrics via visual widgets. The platform distinguishes itself through its data source connectors, which integrate with various SQL, NoSQL, and API-based stores to retrieve information for analysis. It enables self-service analytics by allowing users to run queries with dynamic parameters and supports shared data reporting via public links or embedded dashbo

    Dashboard construction and data visualization for business intelligence.

    Pythonanalyticsathenabi
    在 GitHub 上查看↗28,653
  • pirate/archiveboxpirate 的头像

    pirate/ArchiveBox

    27,721在 GitHub 上查看↗

    ArchiveBox is a self-hosted web archiving system designed to capture and preserve permanent static copies of webpages, media, and PDFs on personal infrastructure. It functions as a digital content curator and personal web archive manager, allowing users to import URLs from bookmarks, RSS feeds, and browser history to create a centralized, searchable knowledge base. The project is distinguished by its ability to archive private, paywalled, or login-protected content using browser cookies and authenticated session persistence. It ensures long-term availability by saving pages in multiple concur

    Self-hosted web archive for creating browsable local backups.

    Python
    在 GitHub 上查看↗27,721
  • iterative/dvciterative 的头像

    iterative/dvc

    15,680在 GitHub 上查看↗

    DVC is a data versioning tool and pipeline orchestrator designed to track large datasets and machine learning models. It functions as a system for managing large data artifacts by storing lightweight metadata in version control while keeping the actual binaries in a separate cache. The project serves as an experiment tracker and remote storage synchronizer, enabling the execution and comparison of machine learning iterations based on hyperparameters and performance metrics. It provides a bridge for pushing and pulling these large data artifacts between local environments and cloud or on-premi

    Version control system for data in machine learning projects.

    Python
    在 GitHub 上查看↗15,680
  • the-paperless-project/paperlessthe-paperless-project 的头像

    the-paperless-project/paperless

    7,917在 GitHub 上查看↗

    Paperless is a self-hosted document management system designed to digitize, index, and archive paper documents. It functions as an optical character recognition system that converts scanned images and PDFs into a searchable digital library, providing a web-based interface for querying and retrieving documents from a database. The system features an automated file ingestion pipeline that monitors specific directories and email inboxes to process and import documents without manual uploading. To maintain a private archive, it includes on-disk encryption for sensitive files and the ability to or

    System for indexing and archiving scanned paper documents.

    Python
    在 GitHub 上查看↗7,917
  • beancount/beancountbeancount 的头像

    beancount/beancount

    5,291在 GitHub 上查看↗

    Beancount is a plain-text double-entry accounting system. It enforces zero-sum transactions, organizes accounts into a hierarchical five-type tree, and verifies balances at specific dates using precision-derived tolerances. Transactions are recorded in plain-text files with a strict syntax that supports currency-specific rounding, automatic interpolation of missing amounts, and comprehensive metadata including tags, links, and payee annotations. Beyond core bookkeeping, Beancount offers investment portfolio tracking with lot-based cost basis management, configurable booking strategies (FIFO,

    Double-entry bookkeeping language for financial records.

    Pythonbeancount
    在 GitHub 上查看↗5,291
  • bram2w/baserowbram2w 的头像

    bram2w/baserow

    5,085在 GitHub 上查看↗

    Baserow 是一个无代码关系数据库和应用程序构建器,允许用户通过可视化界面创建结构化数据表和业务工具。它作为一个无头 REST API 数据后端和自托管数据工作区,提供了一个在保持对数据驻留完全控制的同时管理协作数据库的平台。 该平台集成了大语言模型,作为 LLM 驱动的数据平台,能够从自然语言生成数据库结构、记录内容和技术工作流。它还充当模型上下文协议 (MCP) 服务器,使远程 AI 代理能够以编程方式与结构化数据库记录进行交互。 除了核心数据库功能外,该项目还提供了用于构建品牌外部门户、内部业务应用程序和交互式仪表板的工具。它包括一个用于业务流程自动化的事件驱动自动化引擎,并支持广泛的 API 集成,包括 Webhook、WebSocket 事件流和第三方数据同步。 该软件专为私有基础设施托管和容器化部署而设计,以确保数据主权和安全。

    No-code persistence platform combining databases and spreadsheets.

    Python
    在 GitHub 上查看↗5,085
  • mathesar-foundation/mathesarmathesar-foundation 的头像

    mathesar-foundation/mathesar

    5,012在 GitHub 上查看↗

    Mathesar is a no-code database manager and PostgreSQL GUI that provides a visual interface for managing relational database structures and records. It functions as a low-code data platform for administering schemas, tables, and relationships without the need to write manual SQL commands. The platform allows for the creation of shareable forms to collect data and the management of file attachments linked directly to database records. It includes a PostgreSQL administration tool for controlling database roles, user permissions, and data validation rules. The system covers relational data model

    Spreadsheet-like interface for PostgreSQL databases.

    Svelteairtable-alternativeautomatic-apidatabase-access
    在 GitHub 上查看↗5,012
  • beancount/favabeancount 的头像

    beancount/fava

    2,499在 GitHub 上查看↗

    Fava is a web-based dashboard and query tool for visualizing and analyzing financial records stored in Beancount plain-text ledger files. It serves as a double-entry bookkeeping viewer and plain-text accounting dashboard that renders ledger files as interactive reports, searchable financial tables, and visual tools for exploring balance sheets and income statements. The project distinguishes itself through a specialized BQL query interface that executes SQL-like queries against postings to extract specific financial data and trends. It includes a financial data visualization system for genera

    Web interface for managing double-entry bookkeeping records.

    Pythonbeancountledgerplaintext-accounting
    在 GitHub 上查看↗2,499
  • rd17/ambarRD17 的头像

    RD17/ambar

    1,949在 GitHub 上查看↗

    :mag: Ambar: Document Search Engine

    Document search engine with automated crawling and OCR.

    JavaScript
    在 GitHub 上查看↗1,949
  • camelot-dev/excaliburcamelot-dev 的头像

    camelot-dev/excalibur

    1,790在 GitHub 上查看↗

    Web interface for extracting tabular data from PDF files.

    Pythonextractfor-humanspdf
    在 GitHub 上查看↗1,790
  • artefactual/archivematicaartefactual 的头像

    artefactual/archivematica

    506在 GitHub 上查看↗

    Free and open-source digital preservation system designed to maintain standards-based, long-term access to collections of digital objects.

    Digital preservation system for long-term access to collections.

    Pythonarchivematicadigital-preservation
    在 GitHub 上查看↗506
  1. Home
  2. Part of an Awesome List
  3. Databases & Data
  4. Data Management Systems