awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

Herramientas de catálogo de datos empresariales

Clasificación actualizada el 30 jun 2026

For un catálogo para todos mis datasets, the strongest matches are linkedin/datahub (DataHub is a comprehensive metadata management and data catalog), amundsen-io/amundsen (Amundsen is a fully open-source data catalog and metadata) and open-metadata/openmetadata (OpenMetadata is a fully-featured enterprise data catalog and metadata). datahub-project/datahub and apache/atlas round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.

Plataformas de código abierto para descubrir, catalogar y documentar conjuntos de datos distribuidos en infraestructuras de datos organizacionales complejas.

Herramientas de catálogo de datos empresariales

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • linkedin/datahubAvatar de linkedin

    linkedin/datahub

    12,106Ver en GitHub↗

    DataHub is a metadata management system and data catalog platform designed to provide a centralized directory for discovering, managing, and documenting datasets across a diverse data stack. It serves as a comprehensive framework for metadata management, incorporating a data governance framework to classify sensitive information and assign ownership for organizational accountability. The platform distinguishes itself through AI-enabled data discovery, which connects large language models to a metadata graph to allow for natural language search and exploration of data assets. It also provides

    DataHub is a comprehensive metadata management and data catalog platform that covers automated ingestion, search, lineage, and governance features matching your requirements for a centralized data discovery tool.

    PythonData LineageData Quality ProfilersDatabase Metadata Ingestion
    Ver en GitHub↗12,106
  • amundsen-io/amundsenAvatar de amundsen-io

    amundsen-io/amundsen

    4,737Ver en GitHub↗

    Amundsen is a data catalog and discovery platform that provides a centralized directory for indexing tables and dashboards. It functions as a metadata management system and search engine, allowing users to locate and understand available data assets across diverse distributed sources. The platform includes capabilities for data lineage tracking to map the origin and movement of datasets between systems. It also serves as a data profiling tool, calculating distribution and quality statistics for individual table columns to provide automated insights into the nature of the data. The system man

    Amundsen is a fully open-source data catalog and metadata management platform that provides centralized indexing, search, data lineage tracking, and automated data profiling, matching your need for a self-hosted tool to index, search, annotate, and document datasets for governance and discovery.

    PythonData Quality ProfilersDatabase Metadata IngestionData Discovery Tools
    Ver en GitHub↗4,737
  • open-metadata/openmetadataAvatar de open-metadata

    open-metadata/OpenMetadata

    14,213Ver en GitHub↗

    OpenMetadata is an enterprise data catalog, metadata platform, and governance suite that functions as a knowledge graph for data assets. It serves as an AI-ready metadata layer, providing governed context and organizational memory to large language model agents via the Model Context Protocol. The platform distinguishes itself by capturing institutional knowledge, linking conversations, decisions, and remediation notes directly to data assets to preserve tribal knowledge. It integrates AI agents to automate metadata governance, such as suggesting descriptions and identifying sensitive data thr

    OpenMetadata is a fully-featured enterprise data catalog and metadata platform with automated ingestion, search, data lineage, profiling, quality checks, and collaborative annotation — exactly what you need for centralized data governance and discovery.

    TypeScriptDatabase Metadata IngestionRole-Based Access Control
    Ver en GitHub↗14,213
  • datahub-project/datahubAvatar de datahub-project

    datahub-project/datahub

    12,141Ver en GitHub↗

    DataHub is a metadata management platform designed to unify technical, operational, and business context across diverse data ecosystems. By utilizing a graph-based metadata model and an event-driven ingestion architecture, it creates a centralized source of truth that maps complex data relationships, lineage, and ownership. This foundational framework enables organizations to maintain a synchronized view of their data landscape, supporting both human-led discovery and automated data operations. The platform distinguishes itself through its focus on grounding artificial intelligence and autono

    DataHub is a metadata management platform with graph-based lineage, event-driven automated ingestion, and a centralized source of truth that directly supports data discovery, annotation, and governance across an organization — exactly the kind of comprehensive data catalog this search is after.

    PythonData LineageDatabase Metadata IngestionRole-Based Access Control
    Ver en GitHub↗12,141
  • apache/atlasAvatar de apache

    apache/atlas

    2,110Ver en GitHub↗

    Apache Atlas - Open Metadata Management and Governance capabilities across the Hadoop platform and beyond

    Apache Atlas is a mature, self-hosted metadata management and governance platform that covers automated metadata ingestion, search and discovery, data lineage, collaborative annotation, RBAC, and APIs, making it a comprehensive answer for centralized data cataloging and governance.

    JavaData CatalogsMetadata Management
    Ver en GitHub↗2,110
  • apache/gravitinoAvatar de apache

    apache/gravitino

    2,866Ver en GitHub↗

    Gravitino is a federated metadata lake and unified data catalog designed to manage tables, files, and AI models across diverse data sources and cloud storage. It serves as a centralized interface for governing schemas, access controls, and tagging across relational databases, messaging queues, and object stores. The project distinguishes itself by unifying the management of AI assets, such as machine learning models and their version lineages, alongside traditional tabular data. It also implements the Iceberg REST specification to provide a standardized metadata server and proxy for lakehouse

    Gravitino is a self-hosted, unified data catalog that centralizes metadata governance, search, and access control across databases, files, and AI assets, directly addressing your need for indexing, discovery, and annotation across an organization, though its specific support for collaborative documentation and data profiling is less explicit.

    JavaData LineageRole-Based Access Control
    Ver en GitHub↗2,866
  • ckan/ckanAvatar de ckan

    ckan/ckan

    4,961Ver en GitHub↗

    CKAN is an open-source data management platform that provides the foundation for building data portals. It supports the full lifecycle of datasets—from creation and organization to publishing, cataloging with faceted search, and interactive data visualization—all through a web interface. The platform is built on a modular architecture that includes a plugin-based extensibility system, a harvesting framework for importing metadata from external sources, and a standardized RESTful JSON API for programmatic access to datasets and metadata. The web interface is rendered using the Jinja2 templatin

    CKAN is a mature open-source data portal that lets you catalog, search, and manage datasets with a REST API, harvesting for automated metadata import, and role‑based access—fitting the core of a data catalog, though it lacks built‑in data lineage and profiling.

    PythonData PortalsData CatalogsDataset Catalogs
    Ver en GitHub↗4,961

Related searches

  • una plataforma open source para catálogos de datos
  • sistema de control de versiones para datos de ML
  • toolkit para limpieza y curación de datasets
  • una herramienta de software para gestionar taxonomías jerárquicas
  • an open source tool for collaborative taxonomy
  • un formato columnar compartido en memoria
  • registro para versionado de modelos de ML
  • ver el linaje de mis datos