awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to magda-io/magda

Projects sharing features with Magda

21 open-source projects similar to magda-io/magda, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • linkedin/datahublinkedin avatar

    linkedin/datahub

    12,106View on GitHub↗

    DataHub is a metadata management system and data catalog platform designed to provide a centralized directory for discovering, managing, and documenting datasets across a diverse data stack. It serves as a comprehensive framework for metadata management, incorporating a data governance framework to classify sensitive information and assign ownership for organizational accountability. The platform distinguishes itself through AI-enabled data discovery, which connects large language models to a metadata graph to allow for natural language search and exploration of data assets. It also provides

    Python
    View on GitHub↗12,106
  • ckan/ckanckan avatar

    ckan/ckan

    4,961View on GitHub↗

    CKAN is an open-source data management platform that provides the foundation for building data portals. It supports the full lifecycle of datasets—from creation and organization to publishing, cataloging with faceted search, and interactive data visualization—all through a web interface. The platform is built on a modular architecture that includes a plugin-based extensibility system, a harvesting framework for importing metadata from external sources, and a standardized RESTful JSON API for programmatic access to datasets and metadata. The web interface is rendered using the Jinja2 templatin

    Pythonapicatalogckan
    View on GitHub↗4,961
  • apache/gravitinoapache avatar

    apache/gravitino

    2,866View on GitHub↗

    Gravitino is a federated metadata lake and unified data catalog designed to manage tables, files, and AI models across diverse data sources and cloud storage. It serves as a centralized interface for governing schemas, access controls, and tagging across relational databases, messaging queues, and object stores. The project distinguishes itself by unifying the management of AI assets, such as machine learning models and their version lineages, alongside traditional tabular data. It also implements the Iceberg REST specification to provide a standardized metadata server and proxy for lakehouse

    Javaai-catalogdata-catalogdatalake
    View on GitHub↗2,866

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • apache/hamiltonapache avatar

    apache/hamilton

    2,533View on GitHub↗

    Apache Hamilton — portable & expressive data transformation DAGs

    Jupyter Notebook
    View on GitHub↗2,533
  • apache/incubator-gravitinoA

    apache/incubator-gravitino

    0View on GitHub↗
    View on GitHub↗0
  • ctripcorp/apolloctripcorp avatar

    ctripcorp/apollo

    29,760View on GitHub↗

    Apollo is a microservice configuration management system and dynamic configuration center. It serves as a centralized platform for storing, distributing, and syncing application settings across distributed environments to maintain consistency across various clusters. The system distinguishes itself through a dynamic configuration orchestrator that supports real-time updates to connected applications, eliminating the need for manual service restarts. It features a grayscale configuration deployment tool for rolling out changes to a small subset of service instances and a version control system

    Java
    View on GitHub↗29,760
  • datahub-project/datahubdatahub-project avatar

    datahub-project/datahub

    12,141View on GitHub↗

    DataHub is a metadata management platform designed to unify technical, operational, and business context across diverse data ecosystems. By utilizing a graph-based metadata model and an event-driven ingestion architecture, it creates a centralized source of truth that maps complex data relationships, lineage, and ownership. This foundational framework enables organizations to maintain a synchronized view of their data landscape, supporting both human-led discovery and automated data operations. The platform distinguishes itself through its focus on grounding artificial intelligence and autono

    Pythondata-catalogdata-discoverydata-governance
    View on GitHub↗12,141
  • divanteltd/open-loyaltyD

    DivanteLtd/open-loyalty

    0View on GitHub↗
    View on GitHub↗0
  • elementary-data/elementaryelementary-data avatar

    elementary-data/elementary

    2,373View on GitHub↗

    Elementary OSS: dbt-native data observability

    HTML
    View on GitHub↗2,373
  • grai-io/grai-coregrai-io avatar

    grai-io/grai-core

    315View on GitHub↗
    Python
    View on GitHub↗315
  • amundsen-io/amundsenamundsen-io avatar

    amundsen-io/amundsen

    4,737View on GitHub↗

    Amundsen is a data catalog and discovery platform that provides a centralized directory for indexing tables and dashboards. It functions as a metadata management system and search engine, allowing users to locate and understand available data assets across diverse distributed sources. The platform includes capabilities for data lineage tracking to map the origin and movement of datasets between systems. It also serves as a data profiling tool, calculating distribution and quality statistics for individual table columns to provide automated insights into the nature of the data. The system man

    Pythonamundsendata-catalogdata-discovery
    View on GitHub↗4,737
  • apache/atlasapache avatar

    apache/atlas

    2,110View on GitHub↗

    Apache Atlas - Open Metadata Management and Governance capabilities across the Hadoop platform and beyond

    Java
    View on GitHub↗2,110
  • sodadata/soda-coresodadata avatar

    sodadata/soda-core

    2,292View on GitHub↗
    Pythondata-contractsdata-engineeringdata-governance
    View on GitHub↗2,292
  • marquezproject/marquezMarquezProject avatar

    MarquezProject/marquez

    2,215View on GitHub↗

    Collect, aggregate, and visualize a data ecosystem's metadata

    Java
    View on GitHub↗2,215
  • netflix/genieNetflix avatar

    Netflix/genie

    1,762View on GitHub↗

    Distributed Big Data Orchestration Service

    Java
    View on GitHub↗1,762
  • netflix/metacatNetflix avatar

    Netflix/metacat

    1,684View on GitHub↗

    Metacat is a unified metadata exploration API service. You can explore Hive, RDS, Teradata, Redshift, S3 and Cassandra. Metacat provides you information about what data you have, where it resides and how to process it. Metadata in the end is really data about the data. So the primary purpose of…

    Java
    View on GitHub↗1,684
  • nytimes/gizmonytimes avatar

    nytimes/gizmo

    3,773View on GitHub↗

    Gizmo is a microservice development toolkit and HTTP server framework designed for building distributed services. It provides a collection of libraries for managing service lifecycles, including standardized configuration, logging, and health checks. The toolkit includes a PubSub messaging interface that abstracts the publishing and consuming of messages across different brokers and HTTP endpoints, featuring built-in mocking for integration testing. It also provides a security layer for validating inbound authentication tokens using public key signatures and custom decoders. The project cove

    Gogizmogogoogle-pubsub
    View on GitHub↗3,773
  • odpi/egeriaodpi avatar

    odpi/egeria

    916View on GitHub↗

    Egeria provides the Apache-2.0 licensed open metadata and governance type system, frameworks, APIs, event payloads and interchange protocols to enable tools, engines and platforms to exchange metadata in order to get the best value from data, whilst ensuring it is properly governed.

    Java
    View on GitHub↗916
  • open-metadata/openmetadataopen-metadata avatar

    open-metadata/OpenMetadata

    14,213View on GitHub↗

    OpenMetadata is an enterprise data catalog, metadata platform, and governance suite that functions as a knowledge graph for data assets. It serves as an AI-ready metadata layer, providing governed context and organizational memory to large language model agents via the Model Context Protocol. The platform distinguishes itself by capturing institutional knowledge, linking conversations, decisions, and remediation notes directly to data assets to preserve tribal knowledge. It integrates AI agents to automate metadata governance, such as suggesting descriptions and identifying sensitive data thr

    TypeScriptcontextcontext-layerdata-catalog
    View on GitHub↗14,213
  • opendatadiscovery/odd-platformopendatadiscovery avatar

    opendatadiscovery/odd-platform

    1,410View on GitHub↗

    Next-Gen Data Discovery and Data Observability Platform

    Java
    View on GitHub↗1,410
  • eon01/dockercheatsheeteon01 avatar

    eon01/DockerCheatSheet

    3,938View on GitHub↗

    This project is a comprehensive reference guide and cheat sheet for the Docker CLI. It provides a structured collection of commands and documentation to help users manage container lifecycles, build images, and handle registries. The documentation specifically covers the orchestration of multi-container applications using Docker Compose and the management of scalable services across multiple nodes via Docker Swarm. It also includes detailed guides for configuring virtual networks, bridges, and ports to control container communication. The reference surface extends to container image administ

    View on GitHub↗3,938