9 Repos
Tools and packages for monitoring, testing, and ensuring data integrity.
Explore 9 awesome GitHub repositories matching part of an awesome list · Data Quality. Refine with filters or upvote what's useful.
This project is an exploratory data analysis framework and profiling tool designed to generate comprehensive statistical reports from Pandas and Spark DataFrames. It functions as a data quality profiler that identifies missing values, duplicates, and high correlations within tabular datasets. The tool distinguishes itself through specialized capabilities for time-series analysis, extracting temporal statistics, seasonality, and auto-correlation plots. It also includes a dataset comparison utility to identify structural or content changes between different versions of a dataset. The analysis
Detects missing values, duplicates, and high correlations to ensure data integrity before processing.
Dieses Projekt ist eine Sammlung von Big-Data-Frameworks und Pipelines, darunter ein Apache Hive-Analyse-Framework, eine Plattform für Verhaltensdatenanalyse, eine Predictive-Analytics-Engine und Echtzeit-Datenpipelines. Es bietet die Infrastruktur für den Aufbau von ETL-Workflows (Extract, Transform, Load), um große Datensätze für verteilte Speicherung und SQL-basierte Analysen zu verarbeiten. Das System unterstützt diverse analytische Implementierungen, wie eine Predictive-Engine mittels linearer Regression für Prognosen und eine Echtzeit-Architektur, die Daten über Message-Broker für sofortiges Reporting weiterleitet. Es enthält spezialisierte Funktionen für die Analyse von Nutzerverhalten, E-Commerce-Performance-Messungen und Daten des städtischen Nahverkehrs. Die Codebasis deckt ein breites Spektrum an Data Engineering und Analyse ab, einschließlich Datenbereinigung und -transformation, verteilter Datenaufnahme (Ingestion), fensterbasierter Stream-Verarbeitung und der Visualisierung von Ergebnissen durch Business-Intelligence-Tools. Zudem ermöglicht es die Berechnung spezifischer Geschäftskennzahlen wie Konversionsraten, Monetarisierungs-Performance und Nutzer-Engagement-Level.
Ensures data quality in large datasets by removing duplicate records and standardizing timestamp formats.
Pandera is a data pipeline validation framework and statistical type validation tool. It functions as a library for defining and enforcing schemas on datasets to ensure data quality and consistency, specifically providing validation capabilities for Pandas dataframes. The project includes a schema inference tool that automates setup by analyzing existing dataset samples to generate validation schemas. It also serves as a synthetic data generator, creating artificial datasets based on predefined schemas to verify data-producing functions. The framework covers data engineering quality assuranc
Ensures data integrity and quality by enforcing business logic constraints and validation rules.
Pimcore is an open-source data experience platform that serves as a unified framework for managing product information, digital assets, and customer data. It functions as an enterprise content management system and a master data management platform, providing a centralized source of truth for complex business information. The system is designed to support omnichannel delivery, enabling organizations to publish content and manage digital experiences across diverse platforms through both traditional and headless architectures. The platform distinguishes itself through a metadata-driven object m
Tracks and visualizes data integrity using configurable rules to ensure accuracy across the system.
Elementary OSS: dbt-native data observability
Package for data anomaly detection using dbt tests.
dbt-expectations is an extension package for dbt, inspired by the Great Expectations package for Python. The intent is to allow dbt users to deploy GE-like tests in their data warehouse directly from dbt, vs having to add another integration with their data warehouse.
Port of Great Expectations to extend dbt tests.
Data breaks. Servers break. Your toolchain breaks. Ensure your data team is the first to know and the first to solve with visibility across and down your data estate. Save time with simple, fast data quality test generation and execution. Trust your data, tools, and systems from end to end.
Observability and alerting across data stacks.
The purpose of the dq tool is to make simple storing test results and visualisation of these in a BI dashboard.
Store test results and visualize them in BI dashboards.
A detective for your data. Zero-config data quality monitoring — works with dbt, Postgres, BigQuery, Snowflake. No YAML.
Zero-config CLI for auto-detected data quality anomalies.