For see where my data comes from, the strongest matches are linkedin/datahub (DataHub is a comprehensive open-source metadata platform with automated), open-metadata/openmetadata (OpenMetadata is an enterprise data catalog and governance platform) and datahub-project/datahub (DataHub is a metadata management platform that automatically captures). apache/atlas and cube-js/cube.js round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
These open-source tools map and visualize data movement from raw sources through pipelines to dashboards.
DataHub is a metadata management system and data catalog platform designed to provide a centralized directory for discovering, managing, and documenting datasets across a diverse data stack. It serves as a comprehensive framework for metadata management, incorporating a data governance framework to classify sensitive information and assign ownership for organizational accountability. The platform distinguishes itself through AI-enabled data discovery, which connects large language models to a metadata graph to allow for natural language search and exploration of data assets. It also provides
DataHub is a comprehensive open-source metadata platform with automated column-level lineage extraction, multi-source ingestion, and a rich lineage graph for impact analysis, exactly matching the need for tracking and visualizing data flows across pipelines and BI tools.
OpenMetadata is an enterprise data catalog, metadata platform, and governance suite that functions as a knowledge graph for data assets. It serves as an AI-ready metadata layer, providing governed context and organizational memory to large language model agents via the Model Context Protocol. The platform distinguishes itself by capturing institutional knowledge, linking conversations, decisions, and remediation notes directly to data assets to preserve tribal knowledge. It integrates AI agents to automate metadata governance, such as suggesting descriptions and identifying sensitive data thr
OpenMetadata is an enterprise data catalog and governance platform that natively includes automated lineage capture, column-level lineage visualization, and impact analysis across multiple sources and BI tools, making it a comprehensive open-source solution for tracking data flows and supporting governance.
DataHub is a metadata management platform designed to unify technical, operational, and business context across diverse data ecosystems. By utilizing a graph-based metadata model and an event-driven ingestion architecture, it creates a centralized source of truth that maps complex data relationships, lineage, and ownership. This foundational framework enables organizations to maintain a synchronized view of their data landscape, supporting both human-led discovery and automated data operations. The platform distinguishes itself through its focus on grounding artificial intelligence and autono
DataHub is a metadata management platform that automatically captures and visualizes data lineage across sources, transformations, and dashboards, providing the column-level lineage, impact analysis, and governance features this search requires.
Apache Atlas - Open Metadata Management and Governance capabilities across the Hadoop platform and beyond
Apache Atlas is an open-source metadata management and governance platform that includes automated lineage capture and visualization, enabling impact analysis and governance across Hadoop-based and other sources, fitting the core data lineage need.
Cube is a semantic layer data platform that maps raw SQL databases to standardized business metrics and dimensions. It functions as a SQL dialect translator, converting abstract semantic queries into optimized SQL statements for various cloud data warehouses. The platform operates as a multi-tenant data gateway, isolating information and security permissions for different customers within a single deployment. It includes a relational caching engine that stores pre-aggregated query results to reduce latency and decrease the load on primary data warehouses. The system provides a REST-based int
Cube is a semantic layer platform that standardizes business metrics across databases and exposes them via REST APIs, but it does not automatically capture or visualize data lineage or enable impact analysis across source-to-dashboard flows.
Amundsen is a data catalog and discovery platform that provides a centralized directory for indexing tables and dashboards. It functions as a metadata management system and search engine, allowing users to locate and understand available data assets across diverse distributed sources. The platform includes capabilities for data lineage tracking to map the origin and movement of datasets between systems. It also serves as a data profiling tool, calculating distribution and quality statistics for individual table columns to provide automated insights into the nature of the data. The system man
Amundsen is a data catalog and discovery platform that includes data lineage tracking and visualization, but its primary identity is metadata search and discovery rather than a dedicated lineage observability platform, making it a partial rather than direct fit.
Kedro is a data science pipeline framework and orchestration tool designed to build reproducible and modular data engineering workflows. It functions as an MLOps project template and Python data workflow tool that enforces software engineering best practices to move projects from prototype to production. The system distinguishes itself through a centralized data catalog manager that abstracts data access and versioning across various file formats and cloud storage systems. It further separates processing logic from data access via a lazy-loading data registry and provides a standardized proje
Kedro is a Python pipeline framework that tracks dependencies between pipeline steps, which can serve as a basis for lineage, but it is not a dedicated data lineage or observability platform—it focuses on workflow reproducibility and lacks automated column-level lineage and direct BI tool integration.
sqlglot is a SQL parser and transpiler that represents queries as abstract syntax trees to enable structural analysis, modification, and semantic transformation. It functions as a dialect translator and query optimizer, converting SQL code between different database engines and simplifying syntax trees through rule-based normalization. The project provides a framework for defining custom SQL dialects by overriding tokenizers, parsers, and generators. It includes a lineage analyzer to track data flow from source tables through complex queries to identify the origin of specific columns. Additi
sqlglot is a SQL parser and transpiler with a built-in lineage analyzer that can extract column-level lineage from queries, but it is a library for building lineage into your own system rather than a self-contained data observability platform with visualization, multi-source integration, and metadata management.
JimuReport is an open-source reporting and dashboard engine designed to be embedded directly into Spring Boot applications. Its core identity centers on generating data reports and full-screen dashboards from natural language descriptions, eliminating the need for manual design. The platform also provides a conversational query interface that translates plain-language questions into database queries, returning results as tables and charts without requiring SQL knowledge. What distinguishes JimuReport is its integration of AI skills that can be installed with a single command, enabling report
JimuReport is a reporting and dashboard engine for creating visualisations from databases, but it does not automatically capture data lineage or provide impact analysis across source systems, transformations, and dashboards — it focuses on report generation rather than observability of data flow.
Gravitino is a federated metadata lake and unified data catalog designed to manage tables, files, and AI models across diverse data sources and cloud storage. It serves as a centralized interface for governing schemas, access controls, and tagging across relational databases, messaging queues, and object stores. The project distinguishes itself by unifying the management of AI assets, such as machine learning models and their version lineages, alongside traditional tabular data. It also implements the Iceberg REST specification to provide a standardized metadata server and proxy for lakehouse
Gravitino is a unified metadata catalog and governance layer for diverse data and AI assets, but it does not provide automated capture of end-to-end data flow or lineage visualization for impact analysis, making it a complementary but not standalone lineage platform.
DbGate is a universal database management tool and SQL client that provides a unified interface for querying and administering multiple SQL and NoSQL databases. It functions as a multi-database administration GUI and SQL IDE, allowing users to write and execute scripts and manage database schemas. The project distinguishes itself by acting as an API client and explorer for REST, GraphQL, and OData services, enabling users to fetch and export data from these endpoints. It also serves as a data integration tool, facilitating the movement of records between diverse databases and file formats suc
DbGate is a database management GUI and data integration tool for moving data between systems, but it does not automatically capture column-level lineage, visualize data flows, or provide impact analysis, which are the core capabilities of a data lineage / data observability platform.
OpenBB is a financial data platform and investment research terminal designed to aggregate, normalize, and distribute market data across analytical workflows. It functions as a comprehensive ecosystem that bridges disparate financial data providers with custom applications, spreadsheets, and internal modeling infrastructure. The platform distinguishes itself through a provider-based data abstraction layer that normalizes heterogeneous financial APIs into a consistent, schema-driven format. This architecture supports quantitative research automation and the construction of interactive, widget-
OpenBB is a financial data aggregation and investment research platform, not a tool that automatically tracks and visualises data lineage across source systems to dashboards — it focuses on market data access rather than general-purpose lineage capture and impact analysis.
| Repository | Stars | Language | License | Last push |
|---|---|---|---|---|
| linkedin/datahub | 12.1K | Python | Apache-2.0 | |
| open-metadata/openmetadata | 14.2K | TypeScript | Apache-2.0 | |
| datahub-project/datahub | 12.1K | Python | Apache-2.0 | |
| apache/atlas | 2.1K | Java | Apache-2.0 | |
| cube-js/cube.js | 20.2K | Rust | NOASSERTION | |
| amundsen-io/amundsen | 4.7K | Python | apache-2.0 | |
| kedro-org/kedro | 10.9K | Python | Apache-2.0 | |
| tobymao/sqlglot | 9.3K | Python | MIT | |
| jeecgboot/jimureport | 8.1K | Java | GPL-3.0 | |
| apache/gravitino | 2.9K | Java | apache-2.0 |