awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
open-metadata avatar

open-metadata/OpenMetadata

0
View on GitHub↗
14,213 Stars·2,163 Forks·TypeScript·Apache-2.0·13 Aufrufeopen-metadata.org↗

OpenMetadata

OpenMetadata is an enterprise data catalog, metadata platform, and governance suite that functions as a knowledge graph for data assets. It serves as an AI-ready metadata layer, providing governed context and organizational memory to large language model agents via the Model Context Protocol.

The platform distinguishes itself by capturing institutional knowledge, linking conversations, decisions, and remediation notes directly to data assets to preserve tribal knowledge. It integrates AI agents to automate metadata governance, such as suggesting descriptions and identifying sensitive data through automated stewardship workflows.

The system covers a broad range of data management capabilities, including column-level lineage analysis, data quality observability with freshness and profiling checks, and semantic search for retrieving assets based on conceptual meaning. It also enforces data contracts and compliance policies through a role-based access control framework.

Metadata is collected and synchronized from diverse sources using a pluggable framework of SDKs, APIs, and connectors.

Features

  • Unified Metadata Catalogs - Functions as a unified metadata catalog that aggregates technical metadata, ownership, and quality signals into a searchable graph.
  • Metadata Knowledge Graphs - Stores technical metadata and ownership as a knowledge graph to enable comprehensive lineage tracing and impact analysis.
  • Metadata-to-Agent Bridges - Acts as a metadata-to-agent bridge using the Model Context Protocol to expose governed metadata to LLMs.
  • AI Agent Context Enrichers - Supplies AI agents with business definitions and lineage from the data stack via standard communication protocols.
  • Context Memory Management - Provides a governed history of conversations and remediation notes to maintain context for AI agents.
  • Impact Analysis - Provides column-level lineage and impact analysis to evaluate how changes to data assets affect downstream reports.
  • Data Observability Platforms - Combines data provenance mapping with quality monitoring to analyze the impact of changes.
  • Model Context Protocol Integrations - Implements the Model Context Protocol to enable AI assistants to search and manage metadata via natural language.
  • Semantic Search APIs - Offers semantic search and event-driven APIs to make metadata accessible to external applications and AI agents.
  • Column-Level Lineage Extraction - Traces data flow from source to destination across columns and pipelines to analyze provenance.
  • Conceptual Search Engines - Finds metrics and glossary terms based on conceptual meaning rather than exact keyword matches.
  • Data Asset Discoveries - Locates data through semantic search and descriptions to identify ownership and sample data.
  • Pipeline Metadata Extraction - Automatically captures technical schemas and lineage from orchestration pipelines and technical tools.
  • Data Governance - Enforces data contracts, business glossaries, and ownership policies to maintain a shared business language and compliant data.
  • Data Observability Profilings - Runs profiling checks to detect distribution shifts and triggers alerts for root-cause analysis.
  • Data Quality Contracts - Implements formal agreements between data producers and consumers to enforce schema expectations and quality SLAs.
  • Data Quality Monitors - Ships monitoring for data health, freshness metrics, and profiling checks to detect distribution shifts.
  • Impact Analyzers - Maps relationships between tables and pipelines to visualize data flow and perform downstream impact analysis.
  • AI Grounding Services - Connects AI assistants to governed metadata, ownership, and lineage to ensure factual and trusted response generation.
  • Business Context Grounding - Provides business context grounding to make technical assets discoverable and understandable for users and AI.
  • Database Metadata Ingestion - Automatically extracts and synchronizes metadata from diverse sources using SDKs, APIs, and webhooks.
  • Knowledge Graph Builders - Constructs a knowledge graph connecting technical assets, people, and business concepts.
  • Pluggable Connector Frameworks - Employs a pluggable framework of SDKs and APIs to extract and synchronize metadata from diverse data sources.
  • Context Metadata Governance - Provides governed metadata and organizational memory to AI agents to ensure factual and trusted responses.
  • Semantic Indexing - Indexes metadata based on conceptual meaning and intent to allow retrieval beyond simple keyword matching.
  • Semantic Information Retrieval - Retrieves data assets based on meaning and intent rather than simple keyword matching.
  • Metadata Aggregators - Aggregates technical and operational metadata from the entire ecosystem into a single comprehensive view.
  • Business Glossary Managers - Provides tools for standardizing organizational business definitions and linking them directly to technical data assets.
  • Business Meaning Assignments - Assigns business meaning to raw data structures through the use of glossaries, metrics, and ontologies.
  • Semantic Organizations - Organizes data using glossaries and KPIs to establish a shared business language and lifecycle state.
  • Stewardship Workflows - Provides stewardship workflows to enforce data contracts and classifications, ensuring enterprise-wide data trust and governance.
  • Role-Based Access Control - Manages fine-grained authentication and authorization policies to restrict metadata actions and secure sensitive assets.
  • Provenance Trackers - Captures provenance and pipeline execution metadata using open standards to map data flow between systems.
  • Institutional Knowledge Links - Links unstructured conversations and decisions directly to structured data assets to preserve institutional knowledge.
  • Incident and Thread Context - Preserves context from AI threads and incidents to maintain a historical record of institutional knowledge.
  • Knowledge Management - Captures and preserves domain knowledge by linking descriptive notes to specific assets and workflows.
  • Freshness Monitoring - Monitors freshness signals and metadata to prevent the use of decayed or unreliable data.
  • Data Trust Monitors - Aggregates freshness metrics and certifications to determine if a data asset is trusted for use.
  • Automated Stewardship Workflows - Integrates AI agents to automate metadata governance through description suggestions and sensitive data identification.
  • Governed Agent Memory - Stores conversation logs and AI agent learnings as governed memory for reuse by humans and AI.
  • Decision History Tracking - Preserves historical context by storing remediation notes and decisions associated with data assets.
  • Database Metadata Discovery - Automatically maps and navigates technical asset structures via system catalogs and metadata collection.
  • JSON Schema Modeling - Defines machine-readable asset structures and ontologies using JSON schemas to ensure metadata interoperability.
  • Semantic Search Engines - Implements semantic search to retrieve data assets and business terms based on conceptual meaning and intent.
  • Change Impact Analysis - Evaluates how changes to technical assets affect downstream dependencies to prevent breaking changes in reports.
  • Metadata Workflow Automators - Triggers automated governance actions and webhooks based on changes to the underlying metadata state.
  • Compliance Enforcement Tools - Uses automated systems to verify data integrity and ensure adherence to ownership and compliance policies.
  • Security and Compliance - Combines technical security controls with formal compliance monitoring to ensure assets meet organizational standards.
  • Metadata Normalizations - Standardizes assets and policies into a consistent representation using open schemas and industry standards.
  • Metadata Validations - Applies schemas and shapes to ensure metadata remains consistent across linked data and knowledge graphs.
  • Schema Metadata Definitions - Creates machine-readable metadata definitions using JSON Schemas and ontologies to standardize business context.
  • Metadata Integration APIs - Provides APIs and SDKs to programmatically ingest and update metadata across the software ecosystem.
  • Interoperable Metadata Standards - Applies interoperable schemas and formats to synchronize context between catalogs, knowledge graphs, and external tools.
  • Data Catalogs - Unified metadata management and data discovery platform.
  • Datenmanagement - Platform for data discovery, governance, and lineage tracking.
  • Open Source Catalogs - All-in-one platform for data collaboration, governance, lineage, and quality.
  • Java Projects - Listed in the “Java Projects” section of the Awesome For Beginners awesome list.
  • Python Projects - Listed in the “Python Projects” section of the Awesome For Beginners awesome list.

Star-Verlauf

Star-Verlauf für open-metadata/openmetadataStar-Verlauf für open-metadata/openmetadata

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Häufig gestellte Fragen

Was macht open-metadata/openmetadata?

OpenMetadata is an enterprise data catalog, metadata platform, and governance suite that functions as a knowledge graph for data assets. It serves as an AI-ready metadata layer, providing governed context and organizational memory to large language model agents via the Model Context Protocol.

Was sind die Hauptfunktionen von open-metadata/openmetadata?

Die Hauptfunktionen von open-metadata/openmetadata sind: Unified Metadata Catalogs, Metadata Knowledge Graphs, Metadata-to-Agent Bridges, AI Agent Context Enrichers, Context Memory Management, Impact Analysis, Data Observability Platforms, Model Context Protocol Integrations.

Welche Open-Source-Alternativen gibt es zu open-metadata/openmetadata?

Open-Source-Alternativen zu open-metadata/openmetadata sind unter anderem: datahub-project/datahub — DataHub is a metadata management platform designed to unify technical, operational, and business context across… linkedin/datahub — DataHub is a metadata management system and data catalog platform designed to provide a centralized directory for… dbt-labs/dbt-core — dbt-core is a command-line framework for transforming data within a warehouse using modular SQL and version control.… othmanadi/planning-with-files — Planning with files is an enterprise knowledge graph platform designed to transform unstructured organizational data… opensearch-project/opensearch — OpenSearch is a distributed search and analytics engine designed for indexing, searching, and analyzing massive… unstructured-io/unstructured — Unstructured is an enterprise-grade data orchestration engine designed to transform raw, unstructured files into…

Open-Source-Alternativen zu OpenMetadata

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit OpenMetadata.
  • datahub-project/datahubAvatar von datahub-project

    datahub-project/datahub

    12,141Auf GitHub ansehen↗

    DataHub is a metadata management platform designed to unify technical, operational, and business context across diverse data ecosystems. By utilizing a graph-based metadata model and an event-driven ingestion architecture, it creates a centralized source of truth that maps complex data relationships, lineage, and ownership. This foundational framework enables organizations to maintain a synchronized view of their data landscape, supporting both human-led discovery and automated data operations. The platform distinguishes itself through its focus on grounding artificial intelligence and autono

    Pythondata-catalogdata-discoverydata-governance
    Auf GitHub ansehen↗12,141
  • linkedin/datahubAvatar von linkedin

    linkedin/datahub

    12,106Auf GitHub ansehen↗

    DataHub is a metadata management system and data catalog platform designed to provide a centralized directory for discovering, managing, and documenting datasets across a diverse data stack. It serves as a comprehensive framework for metadata management, incorporating a data governance framework to classify sensitive information and assign ownership for organizational accountability. The platform distinguishes itself through AI-enabled data discovery, which connects large language models to a metadata graph to allow for natural language search and exploration of data assets. It also provides

    Python
    Auf GitHub ansehen↗12,106
  • dbt-labs/dbt-coreAvatar von dbt-labs

    dbt-labs/dbt-core

    13,051Auf GitHub ansehen↗

    dbt-core is a command-line framework for transforming data within a warehouse using modular SQL and version control. It functions as a data transformation engine that enables users to define data structures and business logic through declarative configuration files, which the system then compiles into executable code. By managing complex data dependencies through a directed acyclic graph, it ensures that transformation tasks execute in the correct order while maintaining a manifest-driven state to track lineage and execution history. The project distinguishes itself through an adapter-based d

    Rustanalyticsbusiness-intelligencedata-modeling
    Auf GitHub ansehen↗13,051
  • othmanadi/planning-with-filesAvatar von OthmanAdi

    OthmanAdi/planning-with-files

    14,139Auf GitHub ansehen↗

    Planning with files is an enterprise knowledge graph platform designed to transform unstructured organizational data into a searchable, interconnected network. By utilizing a graph-based retrieval-augmented generation engine, the system grounds language model outputs in verified internal data, ensuring that responses are explainable, traceable, and free from hallucinations. The platform distinguishes itself through a focus on data sovereignty and secure, private infrastructure deployment. It enables organizations to maintain full control over sensitive information by processing data locally o

    Pythonadalagentagent-skills
    Auf GitHub ansehen↗14,139
Alle 30 Alternativen zu OpenMetadata anzeigen→