awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

9 repository-uri

Awesome GitHub RepositoriesData Quality

Tools and packages for monitoring, testing, and ensuring data integrity.

Explore 9 awesome GitHub repositories matching part of an awesome list · Data Quality. Refine with filters or upvote what's useful.

Awesome Data Quality GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • ydataai/pandas-profilingAvatar ydataai

    ydataai/pandas-profiling

    13,610Vezi pe GitHub↗

    This project is an exploratory data analysis framework and profiling tool designed to generate comprehensive statistical reports from Pandas and Spark DataFrames. It functions as a data quality profiler that identifies missing values, duplicates, and high correlations within tabular datasets. The tool distinguishes itself through specialized capabilities for time-series analysis, extracting temporal statistics, seasonality, and auto-correlation plots. It also includes a dataset comparison utility to identify structural or content changes between different versions of a dataset. The analysis

    Detects missing values, duplicates, and high correlations to ensure data integrity before processing.

    Python
    Vezi pe GitHub↗13,610
  • turboway/bigdata_analyseAvatar TurboWay

    TurboWay/bigdata_analyse

    5,238Vezi pe GitHub↗

    Acest proiect este o colecție de framework-uri și pipeline-uri de big data, incluzând un framework de analiză Apache Hive, o platformă de analiză a datelor comportamentale, un motor de analiză predictivă și pipeline-uri de date în timp real. Oferă infrastructura necesară pentru construirea fluxurilor de lucru ETL (Extract, Transform, Load) pentru procesarea seturilor mari de date în vederea stocării distribuite și a analizei bazate pe SQL. Sistemul suportă implementări analitice diverse, cum ar fi un motor predictiv care utilizează regresia liniară pentru prognoza valorilor și o arhitectură în timp real care transmite datele prin message broker-e pentru raportare imediată. Include capabilități specializate pentru analiza comportamentului utilizatorilor, măsurarea performanței în e-commerce și analiza datelor de tranzit urban. Codul sursă acoperă o arie largă de inginerie și analiză a datelor, inclusiv curățarea și transformarea datelor, ingestia distribuită, procesarea fluxurilor bazată pe ferestre (window-based) și vizualizarea rezultatelor prin instrumente de business intelligence. De asemenea, permite calcularea unor metrici de business specifice, cum ar fi ratele de conversie, performanța monetizării și nivelurile de implicare a utilizatorilor.

    Ensures data quality in large datasets by removing duplicate records and standardizing timestamp formats.

    Pythonhqlpythonsql
    Vezi pe GitHub↗5,238
  • unionai-oss/panderaAvatar unionai-oss

    unionai-oss/pandera

    4,382Vezi pe GitHub↗

    Pandera is a data pipeline validation framework and statistical type validation tool. It functions as a library for defining and enforcing schemas on datasets to ensure data quality and consistency, specifically providing validation capabilities for Pandas dataframes. The project includes a schema inference tool that automates setup by analyzing existing dataset samples to generate validation schemas. It also serves as a synthetic data generator, creating artificial datasets based on predefined schemas to verify data-producing functions. The framework covers data engineering quality assuranc

    Ensures data integrity and quality by enforcing business logic constraints and validation rules.

    Pythonassertionsdata-assertionsdata-check
    Vezi pe GitHub↗4,382
  • pimcore/pimcoreAvatar pimcore

    pimcore/pimcore

    3,784Vezi pe GitHub↗

    Pimcore is an open-source data experience platform that serves as a unified framework for managing product information, digital assets, and customer data. It functions as an enterprise content management system and a master data management platform, providing a centralized source of truth for complex business information. The system is designed to support omnichannel delivery, enabling organizations to publish content and manage digital experiences across diverse platforms through both traditional and headless architectures. The platform distinguishes itself through a metadata-driven object m

    Tracks and visualizes data integrity using configurable rules to ensure accuracy across the system.

    PHP
    Vezi pe GitHub↗3,784
  • elementary-data/elementaryAvatar elementary-data

    elementary-data/elementary

    2,373Vezi pe GitHub↗

    Elementary OSS: dbt-native data observability

    Package for data anomaly detection using dbt tests.

    HTML
    Vezi pe GitHub↗2,373
  • calogica/dbt-expectationsAvatar calogica

    calogica/dbt-expectations

    1,228Vezi pe GitHub↗

    dbt-expectations is an extension package for dbt, inspired by the Great Expectations package for Python. The intent is to allow dbt users to deploy GE-like tests in their data warehouse directly from dbt, vs having to add another integration with their data warehouse.

    Port of Great Expectations to extend dbt tests.

    Shell
    Vezi pe GitHub↗1,228
  • datakitchen/data-observability-installerAvatar DataKitchen

    DataKitchen/data-observability-installer

    138Vezi pe GitHub↗

    Data breaks. Servers break. Your toolchain breaks. Ensure your data team is the first to know and the first to solve with visibility across and down your data estate. Save time with simple, fast data quality test generation and execution. Trust your data, tools, and systems from end to end.

    Observability and alerting across data stacks.

    Python
    Vezi pe GitHub↗138
  • infinitelambda/dq-toolsAvatar infinitelambda

    infinitelambda/dq-tools

    54Vezi pe GitHub↗

    The purpose of the dq tool is to make simple storing test results and visualisation of these in a BI dashboard.

    Store test results and visualize them in BI dashboards.

    PLpgSQL
    Vezi pe GitHub↗54
  • rbmuller/scherlokAvatar rbmuller

    rbmuller/scherlok

    6Vezi pe GitHub↗

    A detective for your data. Zero-config data quality monitoring — works with dbt, Postgres, BigQuery, Snowflake. No YAML.

    Zero-config CLI for auto-detected data quality anomalies.

    Python
    Vezi pe GitHub↗6
  1. Home
  2. Part of an Awesome List
  3. Databases & Data
  4. Data Quality