awesome-repositories.com
ब्लॉग
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेसMCP सर्वर
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
linkedin avatar

linkedin/datahub

0
View on GitHub↗
12,106 स्टार्स·3,516 फोर्क्स·Python·Apache-2.0·11 व्यूज़datahub.com↗

Datahub

DataHub is a metadata management system and data catalog platform designed to provide a centralized directory for discovering, managing, and documenting datasets across a diverse data stack. It serves as a comprehensive framework for metadata management, incorporating a data governance framework to classify sensitive information and assign ownership for organizational accountability.

The platform distinguishes itself through AI-enabled data discovery, which connects large language models to a metadata graph to allow for natural language search and exploration of data assets. It also provides specialized data lineage tools that map column-level dependencies to track the flow of data from source to consumption.

The system covers a broad range of capabilities including universal metadata search, data quality monitoring for schema drift and freshness, and dataset profiling. It utilizes a plugin-based ingestion framework to automate the extraction of schemas and usage metrics from warehouses and business intelligence tools.

Features

  • Data Catalogs - Acts as a centralized directory for discovering, managing, and governing diverse organizational data assets.
  • Data Catalogs - Serves as a centralized directory for discovering, managing, and documenting datasets and metadata across a diverse data stack.
  • Metadata-to-Agent Bridges - Provides a protocol that bridges the metadata graph to AI agents for natural language data discovery.
  • Data Lineage - Tracks the origin and transformation history of data assets to understand upstream and downstream dependencies.
  • Natural Language Data Exploration - Uses NLP and metadata graphs to generate insights and explore datasets through natural language.
  • Column-Level Lineage Extraction - Maps granular column-level dependencies to track the flow of data from source to consumption.
  • AI-Powered Exploration - Integrates large language models with a metadata graph to enable natural language search and discovery.
  • Data Governance - Provides frameworks for managing data policies, classifying sensitive info, and assigning organizational ownership.
  • Database Metadata Ingestion - Automates the extraction of technical and operational metadata from external warehouses and BI tools.
  • Enterprise Metadata Management - Provides a platform for ingesting and querying technical and operational metadata from warehouses and BI tools.
  • Metadata Knowledge Graphs - Stores data assets and their relationships as a directed graph to facilitate complex lineage tracking.
  • Metadata Search Engines - Implements a search interface to index and query structured metadata for precise retrieval of datasets across the data stack.
  • Data Asset Lifecycle Management - Tracks data assets and their ownership throughout their entire lifecycle to ensure organizational oversight.
  • Pipeline Metadata Extraction - Automates the extraction of schemas and usage metrics from warehouses and BI tools.
  • Data Quality Monitors - Provides a tracking system for monitoring data freshness, volume changes, and schema drift to ensure ecosystem reliability.
  • Data Quality Profilers - Generates metadata profiles including schemas, data statistics, and technical documentation for individual datasets.
  • Metadata Indexing - Utilizes Elasticsearch to index metadata entities for high-performance full-text search and filtering.
  • Metadata Schema Extensions - Implements a flexible metadata structure that allows custom entity types and attributes without a rigid schema.
  • Metadata Querying - Offers programmatic interfaces and SDKs for retrieving and updating catalog information and custom properties.
  • REST APIs - Provides a standardized REST API for programmatically managing catalog entities and properties.
  • Data Ingestion Plugins - Employs a modular plugin framework of connectors to ingest metadata from various external sources.
  • Metadata Event Processors - Processes data asset changes via an asynchronous event stream to keep the metadata graph consistent.
  • Data Catalog - Metadata search and discovery tool.
  • Data Exchange and ETL - Metadata platform for data stacks.
  • Repository and Developer Analysis - Generalized metadata search and discovery for organizational data.

स्टार हिस्ट्री

linkedin/datahub के लिए स्टार हिस्ट्री चार्टlinkedin/datahub के लिए स्टार हिस्ट्री चार्ट

AI सर्च

और अधिक बेहतरीन रिपॉजिटरी खोजें

अपनी ज़रूरत को सरल भाषा में बताएं — AI हजारों क्यूरेटेड ओपन-सोर्स प्रोजेक्ट्स को प्रासंगिकता के आधार पर रैंक करता है।

Start searching with AI

Datahub के ओपन-सोर्स विकल्प

समान ओपन-सोर्स प्रोजेक्ट्स, जो Datahub के साथ साझा की गई सुविधाओं के आधार पर रैंक किए गए हैं।
  • datahub-project/datahubdatahub-project का अवतार

    datahub-project/datahub

    12,141GitHub पर देखें↗

    DataHub is a metadata management platform designed to unify technical, operational, and business context across diverse data ecosystems. By utilizing a graph-based metadata model and an event-driven ingestion architecture, it creates a centralized source of truth that maps complex data relationships, lineage, and ownership. This foundational framework enables organizations to maintain a synchronized view of their data landscape, supporting both human-led discovery and automated data operations. The platform distinguishes itself through its focus on grounding artificial intelligence and autono

    Pythondata-catalogdata-discoverydata-governance
    GitHub पर देखें↗12,141
  • open-metadata/openmetadataopen-metadata का अवतार

    open-metadata/OpenMetadata

    14,213GitHub पर देखें↗

    OpenMetadata is an enterprise data catalog, metadata platform, and governance suite that functions as a knowledge graph for data assets. It serves as an AI-ready metadata layer, providing governed context and organizational memory to large language model agents via the Model Context Protocol. The platform distinguishes itself by capturing institutional knowledge, linking conversations, decisions, and remediation notes directly to data assets to preserve tribal knowledge. It integrates AI agents to automate metadata governance, such as suggesting descriptions and identifying sensitive data thr

    TypeScriptcontextcontext-layerdata-catalog
    GitHub पर देखें↗14,213
  • amundsen-io/amundsenamundsen-io का अवतार

    amundsen-io/amundsen

    4,737GitHub पर देखें↗

    Amundsen is a data catalog and discovery platform that provides a centralized directory for indexing tables and dashboards. It functions as a metadata management system and search engine, allowing users to locate and understand available data assets across diverse distributed sources. The platform includes capabilities for data lineage tracking to map the origin and movement of datasets between systems. It also serves as a data profiling tool, calculating distribution and quality statistics for individual table columns to provide automated insights into the nature of the data. The system man

    Pythonamundsendata-catalogdata-discovery
    GitHub पर देखें↗4,737
  • dbt-labs/dbt-coredbt-labs का अवतार

    dbt-labs/dbt-core

    13,051GitHub पर देखें↗

    dbt-core is a command-line framework for transforming data within a warehouse using modular SQL and version control. It functions as a data transformation engine that enables users to define data structures and business logic through declarative configuration files, which the system then compiles into executable code. By managing complex data dependencies through a directed acyclic graph, it ensures that transformation tasks execute in the correct order while maintaining a manifest-driven state to track lineage and execution history. The project distinguishes itself through an adapter-based d

    Rustanalyticsbusiness-intelligencedata-modeling
    GitHub पर देखें↗13,051
Datahub के सभी 30 विकल्प देखें→

अक्सर पूछे जाने वाले प्रश्न

linkedin/datahub क्या करता है?

DataHub is a metadata management system and data catalog platform designed to provide a centralized directory for discovering, managing, and documenting datasets across a diverse data stack. It serves as a comprehensive framework for metadata management, incorporating a data governance framework to classify sensitive information and assign ownership for organizational accountability.

linkedin/datahub की मुख्य विशेषताएं क्या हैं?

linkedin/datahub की मुख्य विशेषताएं हैं: Data Catalogs, Metadata-to-Agent Bridges, Data Lineage, Natural Language Data Exploration, Column-Level Lineage Extraction, AI-Powered Exploration, Data Governance, Database Metadata Ingestion।

linkedin/datahub के कुछ ओपन-सोर्स विकल्प क्या हैं?

linkedin/datahub के ओपन-सोर्स विकल्पों में शामिल हैं: datahub-project/datahub — DataHub is a metadata management platform designed to unify technical, operational, and business context across… open-metadata/openmetadata — OpenMetadata is an enterprise data catalog, metadata platform, and governance suite that functions as a knowledge… amundsen-io/amundsen — Amundsen is a data catalog and discovery platform that provides a centralized directory for indexing tables and… dbt-labs/dbt-core — dbt-core is a command-line framework for transforming data within a warehouse using modular SQL and version control.… ckan/ckan — CKAN is an open-source data management platform that provides the foundation for building data portals. It supports… apache/gravitino — Gravitino is a federated metadata lake and unified data catalog designed to manage tables, files, and AI models across…