awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
amundsen-io avatar

amundsen-io/amundsen

0
View on GitHub↗

Amundsen

Amundsen is a data catalog and discovery platform that provides a centralized directory for indexing tables and dashboards. It functions as a metadata management system and search engine, allowing users to locate and understand available data assets across diverse distributed sources.

The platform includes capabilities for data lineage tracking to map the origin and movement of datasets between systems. It also serves as a data profiling tool, calculating distribution and quality statistics for individual table columns to provide automated insights into the nature of the data.

The system manages technical metadata ingestion from various databases and orchestration tools through a metadata integration workflow. These capabilities are supported by a search index for data discovery, a relational metadata store, and a REST-based API for external tool integration.

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI
www.amundsen.io/amundsen
↗

Features

  • Data Catalogs - Provides a centralized directory for indexing tables and dashboards to help users discover and understand organizational data assets.
  • Search and Indexing - Indexes tables and dashboards to create a searchable interface for discovering available data assets.
  • Data Asset Search - Provides a specialized search engine for indexing and retrieving distributed data assets across an organization.
  • Data Asset Discoveries - Locates specific tables and dashboards across an organization using a centralized searchable index.
  • Data Discovery Tools - Implements a searchable interface and index for locating specific datasets across diverse distributed data sources.
  • Multi-Source Data Integration - Connects to various databases and orchestration tools to ingest metadata into a central location.
  • Metadata Sync Engines - Synchronizes technical schemas and metadata from diverse databases and orchestration tools into a central repository.
  • Table Metadata Inspection - Allows users to visualize table schemas and column details to understand the structural organization of datasets.
  • Database Metadata Ingestion - Automates the extraction and synchronization of technical metadata from various database instances.
  • Plugin-Based Ingestion - Uses a modular plugin interface to ingest technical metadata from various external databases and tools.
  • Metadata Indexing - Indexes technical metadata into a search engine to enable fast discovery of datasets across distributed sources.
  • Metadata Management Systems - Manages the ingestion and visualization of schemas and column-level statistics to ensure data quality and lineage.
  • Asset Lineage Visualizations - Tracks and visualizes the flow of data between different systems to map the origin and movement of datasets.
  • Data Lineage Trackers - Maps the origin and movement of datasets between systems to track data provenance.
  • Data Quality Profilers - Calculates distribution and quality statistics for table columns to provide automated data quality insights.
  • Single-Pass Column Analyzers - Computes descriptive statistics and value frequencies for table columns to reveal data quality insights.
  • Metadata REST Endpoints - Exposes the metadata store through REST endpoints for frontend consumption and external tool integration.
  • Column-Level Validation - Generates distribution and quality statistics for individual table columns to understand dataset nature.
  • Data Catalogs - Metadata-driven data discovery and exploration engine.
  • Metadata Management - Metadata-driven application for data analyst productivity.
4,737 estrellas·970 forks·Python·apache-2.0·24 vistas

Historial de estrellas

Gráfico del historial de estrellas de amundsen-io/amundsenGráfico del historial de estrellas de amundsen-io/amundsen

Alternativas open-source a Amundsen

Proyectos open-source similares, clasificados según cuántas características comparten con Amundsen.
  • datahub-project/datahubAvatar de datahub-project

    datahub-project/datahub

    12,141Ver en GitHub↗

    DataHub is a metadata management platform designed to unify technical, operational, and business context across diverse data ecosystems. By utilizing a graph-based metadata model and an event-driven ingestion architecture, it creates a centralized source of truth that maps complex data relationships, lineage, and ownership. This foundational framework enables organizations to maintain a synchronized view of their data landscape, supporting both human-led discovery and automated data operations. The platform distinguishes itself through its focus on grounding artificial intelligence and autono

    Pythondata-catalogdata-discoverydata-governance
    Ver en GitHub↗12,141
  • linkedin/datahubAvatar de linkedin

    linkedin/datahub

    12,106Ver en GitHub↗

    DataHub is a metadata management system and data catalog platform designed to provide a centralized directory for discovering, managing, and documenting datasets across a diverse data stack. It serves as a comprehensive framework for metadata management, incorporating a data governance framework to classify sensitive information and assign ownership for organizational accountability. The platform distinguishes itself through AI-enabled data discovery, which connects large language models to a metadata graph to allow for natural language search and exploration of data assets. It also provides

    Python
    Ver en GitHub↗12,106
  • open-metadata/openmetadataAvatar de open-metadata

    open-metadata/OpenMetadata

    14,213Ver en GitHub↗

    OpenMetadata is an enterprise data catalog, metadata platform, and governance suite that functions as a knowledge graph for data assets. It serves as an AI-ready metadata layer, providing governed context and organizational memory to large language model agents via the Model Context Protocol. The platform distinguishes itself by capturing institutional knowledge, linking conversations, decisions, and remediation notes directly to data assets to preserve tribal knowledge. It integrates AI agents to automate metadata governance, such as suggesting descriptions and identifying sensitive data thr

    TypeScriptcontextcontext-layerdata-catalog
    Ver en GitHub↗14,213
  • awslabs/aws-data-wranglerAvatar de awslabs

    awslabs/aws-data-wrangler

    4,107Ver en GitHub↗

    This project is an AWS pandas integration library and data pipeline framework designed to simplify the movement and transformation of data between local memory and AWS storage and analytics services. It functions as a cloud data lake toolkit and storage file manager, allowing users to read, write, and transform structured data across various cloud environments. The library distinguishes itself as a distributed compute orchestrator capable of managing clusters in environments such as EMR to process datasets that exceed the memory limits of a single machine. It also provides specialized capabil

    Python
    Ver en GitHub↗4,107
Ver las 30 alternativas a Amundsen→

Preguntas frecuentes

¿Qué hace amundsen-io/amundsen?

Amundsen is a data catalog and discovery platform that provides a centralized directory for indexing tables and dashboards. It functions as a metadata management system and search engine, allowing users to locate and understand available data assets across diverse distributed sources.

¿Cuáles son las características principales de amundsen-io/amundsen?

Las características principales de amundsen-io/amundsen son: Data Catalogs, Search and Indexing, Data Asset Search, Data Asset Discoveries, Data Discovery Tools, Multi-Source Data Integration, Metadata Sync Engines, Table Metadata Inspection.

¿Qué alternativas de código abierto existen para amundsen-io/amundsen?

Las alternativas de código abierto para amundsen-io/amundsen incluyen: datahub-project/datahub — DataHub is a metadata management platform designed to unify technical, operational, and business context across… linkedin/datahub — DataHub is a metadata management system and data catalog platform designed to provide a centralized directory for… open-metadata/openmetadata — OpenMetadata is an enterprise data catalog, metadata platform, and governance suite that functions as a knowledge… awslabs/aws-data-wrangler — This project is an AWS pandas integration library and data pipeline framework designed to simplify the movement and… apache/gravitino — Gravitino is a federated metadata lake and unified data catalog designed to manage tables, files, and AI models across… ckan/ckan — CKAN is an open-source data management platform that provides the foundation for building data portals. It supports…