awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

461 个仓库

Awesome GitHub RepositoriesData Governance and Modeling

Frameworks for defining schemas, ensuring standardization, and managing data assets and sovereignty.

Explore 461 awesome GitHub repositories matching data & databases · Data Governance and Modeling. Refine with filters or upvote what's useful.

Awesome Data Governance and Modeling GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • vinta/awesome-pythonvinta 的头像

    vinta/awesome-python

    303,207在 GitHub 上查看↗

    这是一个全面的、由社区策划的目录,组织了庞大的 Python 软件库、框架和工具生态。它作为一个中心化知识库,旨在促进生态导航并加速开发者在整个软件开发生命周期中的发现过程。 该目录通过提供按技术领域分类的结构化资源索引脱颖而出,范围从基础开发工具到专业工程领域。它涵盖了人工智能、数据科学、Web 开发和基础设施管理等高级能力,使开发者能够为特定的技术挑战识别经过验证的解决方案。 该项目涵盖了广泛的能力领域,包括依赖管理、静态代码分析和自动化测试工具。它还编目了用于持久数据存储、云基础设施编排和接口开发的资源,为构建和维护复杂软件系统提供了统一的参考。

    Verify data integrity by applying schema constraints and type requirements to incoming information.

    Pythonawesomecollectionspython
    在 GitHub 上查看↗303,207
  • ossu/computer-scienceossu 的头像

    ossu/computer-science

    205,190在 GitHub 上查看↗

    本项目提供了一个为自学者设计的结构化计算机科学课程框架。它将开放获取的学术资源(包括教科书、讲座和作业)组织成一条与正式本科学位要求相呼应的连贯路径。通过将理论学习与实际软件工程方法论相结合,该平台使学生能够独立掌握基础概念和高级技术技能。 该课程的独特之处在于利用基于版本控制的工作流来管理教育体验。学习者使用基于仓库的工具来跟踪学术里程碑、维护已完成作业的持久历史记录,并根据既定要求验证其技术解决方案。这种方法鼓励在学习过程中采用行业标准的工程实践,例如配置隔离的开发环境和管理项目依赖项。 该平台支持广泛的技术开发,涵盖计算问题解决、面向对象设计和数据分析等领域。它通过社区驱动的平台促进协作学习,使学生能够进行同行互动并验证彼此的工作。该课程作为开源资源进行维护,为构建软件工程的专业能力提供了全面的指南。

    Verify text formats within incoming data streams by applying regular expressions to ensure that all information matches expected structures before processing continues further.

    HTMLawesome-listcomputer-sciencecourses
    在 GitHub 上查看↗205,190
  • avelino/awesome-goavelino 的头像

    avelino/awesome-go

    175,576在 GitHub 上查看↗

    This project serves as a comprehensive language ecosystem index, functioning as a centralized, community-curated directory for the Go programming language. It organizes a vast landscape of software components, libraries, and development tools into a structured, navigable hierarchy, enabling developers to efficiently discover resources tailored to specific functional domains. The repository distinguishes itself through a decentralized contribution model, where community-driven updates ensure the index remains current with the rapidly evolving software landscape. Beyond simple resource listing,

    Identifies in-memory data stores and distributed caching solutions that support record expiration.

    Goawesomeawesome-listgo
    在 GitHub 上查看↗175,576
  • snailclimb/javaguideSnailclimb 的头像

    Snailclimb/JavaGuide

    156,541在 GitHub 上查看↗

    This project is a comprehensive educational repository providing technical documentation and learning materials across a wide range of computer science and software engineering domains. It serves as a centralized knowledge base for developers, covering core programming concepts, database management, distributed systems, and system design principles. The content spans fundamental Java programming, including collection frameworks and runtime environments, alongside deep dives into web communication protocols and browser internals. It also provides extensive resources on database internals, such

    Demonstrates architectural approaches for generating unique identifiers to ensure data integrity across partitioned database shards.

    JavaScriptalgorithmsdistributed-systemsinterview
    在 GitHub 上查看↗156,541
  • logspace-ai/langflowlogspace-ai 的头像

    logspace-ai/langflow

    149,776在 GitHub 上查看↗

    Langflow is a low-code platform for designing and deploying multi-step AI agent pipelines and large language model sequences. It provides a visual environment to map logic and data flow between components, serving as an orchestrator for managing conversations and data retrieval across multiple autonomous agents. The platform distinguishes itself through a drag-and-drop interface that allows for the construction of complex AI pipelines without extensive boilerplate code. It enables the conversion of these internal workflows into standardized tools for external connectivity via the Model Contex

    Syncs the visual drag-and-drop interface with a JSON representation of workflow logic for consistent state management.

    Python
    在 GitHub 上查看↗149,776
  • langchain-ai/langchainlangchain-ai 的头像

    langchain-ai/langchain

    139,458在 GitHub 上查看↗

    LangChain is an orchestration framework designed for building, managing, and deploying applications powered by large language models. It provides a unified integration layer that normalizes disparate model provider APIs into a consistent set of primitives, enabling developers to build complex, multi-step AI workflows that manage state, memory, and tool execution. The project distinguishes itself through a durable execution runtime that maintains persistent state across long-running processes by checkpointing progress to external storage. It models agent workflows as directed graphs, allowing

    Configure data paths to specify storage locations for session logs, configuration files, and agent customizations.

    Pythonagentsaiai-agents
    在 GitHub 上查看↗139,458
  • chalarangelo/30-seconds-of-codeChalarangelo 的头像

    Chalarangelo/30-seconds-of-code

    128,121在 GitHub 上查看↗

    30-seconds-of-code is a comprehensive knowledge base and programming snippet library designed to support software engineering education and professional development. It provides a curated collection of reusable code units and technical guides that help developers master core language mechanics, design patterns, and architectural philosophies. The project distinguishes itself by offering a wide-ranging library of algorithmic solutions and web development patterns that are organized into modular, independently testable units. It emphasizes functional programming paradigms and declarative logic,

    Provides frameworks for enforcing data constraints and integrity rules.

    JavaScriptastroawesome-listcss
    在 GitHub 上查看↗128,121
  • iptv-org/iptviptv-org 的头像

    iptv-org/iptv

    127,909在 GitHub 上查看↗

    This project is a community-maintained, open-source repository that functions as a centralized directory for streaming metadata. It aggregates publicly available network stream links and organizes them into standardized, machine-readable playlist formats. By acting strictly as a metadata-only index, the platform enables users to access and organize live broadcast content across various third-party media playback applications without hosting or distributing any actual video files. The repository distinguishes itself through a collaborative, crowdsourced workflow where contributors actively mai

    Stores and manages a centralized directory of streaming channel metadata for reliable external data access.

    TypeScriptiptvm3uplaylist
    在 GitHub 上查看↗127,909
  • ripienaar/free-for-devripienaar 的头像

    ripienaar/free-for-dev

    123,154在 GitHub 上查看↗

    This project is a community-maintained directory of technical resources, tools, and services that offer free tiers for developers. It serves as a centralized reference point for discovering infrastructure, software, and educational materials, helping individuals and teams minimize operational costs while building and scaling applications. The directory distinguishes itself through a collaborative, community-driven curation model that aggregates metadata about third-party services. By utilizing a hierarchical taxonomy and storing all content in version-controlled, plain-text files, the project

    Sorts technical resources into a structured taxonomy to improve searchability and discovery across specialized development domains.

    HTMLawesome-listfree-for-developers
    在 GitHub 上查看↗123,154
  • immich-app/immichimmich-app 的头像

    immich-app/immich

    104,236在 GitHub 上查看↗

    Immich is a self-hosted media management platform designed to provide a centralized, private repository for photos and videos. It functions as a comprehensive system for organizing, backing up, and viewing personal media collections across mobile devices, web browsers, and external storage locations. By maintaining full control over data ownership and storage infrastructure, the platform ensures that users retain sovereignty over their digital assets. The system distinguishes itself through a distributed architecture that coordinates background media synchronization, real-time filesystem moni

    Archives critical media and user-specific files to ensure complete recovery alongside database metadata.

    TypeScriptbackup-toolfluttergoogle-photos
    在 GitHub 上查看↗104,236
  • tiangolo/fastapitiangolo 的头像

    tiangolo/fastapi

    99,301在 GitHub 上查看↗

    FastAPI is a high-performance Python web framework designed for building REST APIs. It operates as an ASGI web framework, providing a system to create structured HTTP endpoints that automatically serialize data and validate request parameters. The framework utilizes Python type hints to drive data validation and serialization, automatically generating machine-readable OpenAPI and JSON Schema specifications. This process enables the automatic creation of interactive, browser-based API documentation where endpoints can be tested directly. The project includes a dependency injection system for

    Validates incoming HTTP request bodies, parameters, and headers against defined types with automatic error handling.

    Python
    在 GitHub 上查看↗99,301
  • gin-gonic/gingin-gonic 的头像

    gin-gonic/gin

    88,694在 GitHub 上查看↗

    Gin is a web framework designed for building high-performance web services and APIs. It functions as a middleware-oriented engine that processes incoming HTTP requests through a sequential chain of handlers, allowing for the modular management of cross-cutting concerns such as authentication and logging. The framework utilizes a radix tree data structure to perform request routing, ensuring high-speed path matching with minimal memory overhead. It distinguishes itself by employing a zero-reflection dispatch mechanism that invokes handler functions through static type assertions, avoiding the

    Enforces data integrity by mapping incoming payloads directly into structured models while applying strict validation rules.

    Goframeworkgingo
    在 GitHub 上查看↗88,694
  • home-assistant/corehome-assistant 的头像

    home-assistant/core

    87,753在 GitHub 上查看↗

    Home Assistant is a centralized home automation platform designed to orchestrate diverse internet-connected devices and services. It functions as a local-first control system that normalizes heterogeneous hardware protocols into a unified set of entities, attributes, and services. The core architecture relies on an event-driven state bus and a modular integration model, allowing the system to manage state changes and communicate across decoupled components through standardized interfaces. The platform distinguishes itself through a highly flexible, declarative configuration framework that all

    Retains long-term historical entity data, allowing users to analyze trends and rectify inaccurate records.

    Pythonasynciohacktoberfesthome-automation
    在 GitHub 上查看↗87,753
  • syncthing/syncthingsyncthing 的头像

    syncthing/syncthing

    85,400在 GitHub 上查看↗

    Syncthing 是一个去中心化的文件同步引擎,通过点对点网状网络在多个设备间保持一致的数据状态。它作为后台守护进程运行,在受信任的节点之间自动复制文件的创建、修改和删除,无需中央服务器。通过利用内容可寻址块索引和块级增量同步,系统仅识别并传输文件的修改部分,从而确保在异构环境中高效地传播数据。 该项目以安全优先的架构著称,依赖相互 TLS 认证来验证设备身份,确保所有连接在加密上绑定到受信任的证书指纹。它支持灵活的同步模式,包括双向复制、用于备份的单向镜像以及基于引用的强制执行。为了增加隐私性,系统为不受信任的设备提供了文件夹级加密,并允许对网络流量进行细粒度控制,包括限制操作仅在本地网络进行或利用中继基础设施进行 NAT 穿透。 除了核心复制功能外,该平台还提供全面的管理工具,包括用于监控连接状态和吞吐量的 Web 仪表板,以及用于高级配置的命令行界面。它包含强大的版本控制策略以防止数据丢失,并通过原生服务集成和可观测性指标支持复杂的部署场景。该软件专为跨平台兼容性而设计,可通过标准包管理器或容器化环境安装。

    Ensures filesystem consistency by performing atomic renames of temporary files during synchronization.

    Gogop2ppeer-to-peer
    在 GitHub 上查看↗85,400
  • laravel/laravellaravel 的头像

    laravel/laravel

    84,489在 GitHub 上查看↗

    Laravel is a comprehensive full-stack web framework designed for building scalable server-side applications. It provides an integrated development environment that centers on an object-relational mapper for database abstraction, a robust routing system, and a sophisticated service container for dependency injection. The framework is built to handle complex application requirements through a modular architecture that emphasizes convention over configuration. What distinguishes Laravel is its deep integration of background processing and event-driven communication. It features a task queue orch

    Validates data types, formats, and database existence across various input sources using a comprehensive suite of built-in rules.

    Bladeframeworklaravelphp
    在 GitHub 上查看↗84,489
  • infiniflow/ragflowinfiniflow 的头像

    infiniflow/ragflow

    82,922在 GitHub 上查看↗

    This project is a comprehensive retrieval-augmented generation platform designed for building, managing, and deploying knowledge-based AI applications. It provides a unified environment for organizing datasets, configuring conversational chat assistants, and developing autonomous agents that execute multi-step reasoning workflows. By integrating document intelligence with advanced retrieval pipelines, the platform enables the creation of grounded, verifiable responses supported by traceable citations. The platform distinguishes itself through deep document understanding and sophisticated know

    Organizes knowledge by uploading, parsing, and indexing documents into structured datasets for retrieval-augmented generation.

    Pythonagentagenticagentic-ai
    在 GitHub 上查看↗82,922
  • elastic/elasticsearchelastic 的头像

    elastic/elasticsearch

    77,012在 GitHub 上查看↗

    Elasticsearch is a distributed search engine and document store designed for the high-performance indexing and retrieval of massive volumes of unstructured data. It functions as a centralized analytics platform, providing a schema-flexible architecture that organizes information into searchable indices while maintaining global cluster state through a distributed consensus mechanism. The platform distinguishes itself through its integrated approach to observability, security, and advanced analytics. It combines full-text, vector, and hybrid search capabilities with machine learning-driven insi

    Optimizes storage costs by automatically shifting indices between hot, warm, and cold performance tiers based on age and access patterns.

    Javaelasticsearchjavasearch-engine
    在 GitHub 上查看↗77,012
  • redis/redisredis 的头像

    redis/redis

    74,906在 GitHub 上查看↗

    Redis is an in-memory, key-value database designed to provide sub-millisecond latency for read and write operations. It functions as a versatile data platform, serving as a distributed cache, a message broker, a NoSQL document store, and a vector database. The system utilizes an event-driven, single-threaded loop to process requests efficiently, while maintaining data durability through append-only persistence logs and asynchronous snapshotting mechanisms. What distinguishes Redis is its ability to handle complex data structures—including strings, hashes, lists, sets, and sorted sets—alongsid

    Applies time-to-live values to cached entries, preventing memory bloat by automatically evicting stale or expired information.

    Ccachecachingdatabase
    在 GitHub 上查看↗74,906
  • apache/supersetapache 的头像

    apache/superset

    73,451在 GitHub 上查看↗

    Superset is a web-based business intelligence platform designed for data exploration, visualization, and interactive dashboarding. It functions as a query-driven analytics engine that connects to various SQL databases, allowing users to perform ad-hoc analysis, define virtual metrics, and build complex data visualizations through a centralized interface. The platform distinguishes itself through a robust semantic layer that transforms raw database schemas into calculated columns and virtual metrics, enabling consistent business logic across an organization. It features a plugin-based visualiz

    Centralizes governance by applying security policies and metadata management across large-scale organizational data.

    TypeScriptanalyticsapacheapache-superset
    在 GitHub 上查看↗73,451
  • appflowy-io/appflowyAppFlowy-IO 的头像

    AppFlowy-IO/AppFlowy

    72,474在 GitHub 上查看↗

    AppFlowy is a local-first knowledge base and collaborative workspace platform designed for structured information management. It functions as a modular productivity suite where users organize content through a block-based document model, allowing for flexible nesting and granular manipulation of data. The system prioritizes data sovereignty by enabling self-hosted storage, ensuring that sensitive information remains under user control while maintaining offline accessibility. The platform distinguishes itself through a decoupled architecture that separates its high-performance, memory-safe cor

    Prioritizes data privacy by keeping all information stored locally under the user's complete control.

    Dartblogconfluence-alternativecontent-management
    在 GitHub 上查看↗72,474
上一个123456…24下一个
  1. Home
  2. Data & Databases
  3. Data Governance and Modeling

探索子标签

  • Bulk Update ExtensionsTools for performing batch modifications on multiple database records simultaneously.
  • Data Management & Governance9 个子标签Frameworks and policies that ensure data quality, security, compliance, and lifecycle management across an organization.
  • Data Modeling and Schemas11 个子标签Tools and standards used to define, visualize, and evolve the structural organization of data within a system.
  • Data Sovereignty Models1 个子标签Frameworks that enable organizations to maintain control and compliance over data residency and jurisdictional requirements.
  • Data Standardization6 个子标签Utilities that transform and normalize disparate data formats into consistent, standardized structures.
  • System Metadata1 个子标签Systems for managing descriptive information about data entities to improve discoverability and context.
  • Taxonomies1 个子标签Systems for organizing and classifying data into hierarchical or categorical structures for better information retrieval.