awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

12 repository-uri

Awesome GitHub RepositoriesData Quality and Validation

Frameworks for ensuring data integrity and schema compliance.

Explore 12 awesome GitHub repositories matching part of an awesome list · Data Quality and Validation. Refine with filters or upvote what's useful.

Awesome Data Quality and Validation GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • pydantic/pydanticAvatar pydantic

    pydantic/pydantic

    26,932Vezi pe GitHub↗

    Pydantic is a data validation and serialization library that enforces schema constraints and performs type conversion on complex data structures. It utilizes standard Python type annotations to define data models, allowing developers to establish structured schemas that automatically enforce business rules and constraints without the need for custom domain-specific languages. The library distinguishes itself by transforming high-level model definitions into optimized code during initialization to minimize runtime overhead. It supports recursive validation for nested data structures and employ

    Data validation using Python type annotations.

    Pythonhintsjson-schemaparsing
    Vezi pe GitHub↗26,932
  • ludwig-ai/ludwigAvatar ludwig-ai

    ludwig-ai/ludwig

    11,717Vezi pe GitHub↗

    Ludwig is a multimodal machine learning platform and low-code framework designed for building, training, and deploying neural networks. It enables the construction of models that process text, images, audio, and tabular data through a unified interface using declarative configuration files rather than custom code. The system features a specialized low-code framework for large language models, supporting supervised fine-tuning, preference alignment, and a constrained decoding tool to force structured data output via logit extraction. It also includes an automated model architecture search to i

    Checks dataframes for missing values and class imbalances to ensure data consistency before training.

    Pythoncomputer-visiondata-centricdata-science
    Vezi pe GitHub↗11,717
  • yzhao062/pyodAvatar yzhao062

    yzhao062/pyod

    9,878Vezi pe GitHub↗

    PyOD is a Python anomaly detection library used to identify outliers in tabular, time series, graph, text, and image data. It provides a collection of algorithms for detecting anomalous data points and includes a unified detector interface that standardizes input and output signatures across its available detection algorithms. The project features a multi-modal outlier detector for identifying anomalies across diverse formats including unstructured text and images, as well as a specialized toolkit for graph-based and time-series anomaly detection. It includes an ensemble framework for combini

    Outlier and anomaly detection.

    Pythonagentic-aianomaly-detectiondata-mining
    Vezi pe GitHub↗9,878
  • apache/seatunnelAvatar apache

    apache/seatunnel

    9,427Vezi pe GitHub↗

    SeaTunnel is a distributed data integration engine designed to synchronize structured and unstructured data across diverse sources and sinks. It functions as a multi-engine execution framework that can run data integration tasks across different distributed computing backends to optimize workload performance. The project is distinguished by a visual data pipeline designer for configuring workflows without manual code and a specialized change data capture tool for streaming incremental database updates. It also includes an enrichment pipeline that integrates large language models and embedding

    Includes processes for checking records against predefined rules to ensure data integrity during movement.

    Javaapachebatchcdc
    Vezi pe GitHub↗9,427
  • feast-dev/feastAvatar feast-dev

    feast-dev/feast

    6,727Vezi pe GitHub↗

    Feast is an open-source feature store for machine learning that provides a central platform for defining, storing, and serving features across both training and inference workflows. It operates as a declarative system where feature definitions are written as code in Python files, synchronized to a central registry, and made available for low-latency online retrieval or point-in-time correct historical joins for training datasets. The project abstracts storage behind a pluggable architecture, allowing offline and online backends to be swapped without changing retrieval logic, and coordinates ma

    Feast profiles and validates feature data against expectations to catch drift or errors before they affect models.

    Pythonbig-datadata-engineeringdata-quality
    Vezi pe GitHub↗6,727
  • focus-creative-games/lubanAvatar focus-creative-games

    focus-creative-games/luban

    4,458Vezi pe GitHub↗

    Luban este un toolchain de configurare a jocurilor conceput pentru a converti datele bazate pe foi de calcul în formate binare optimizate și cod sursă type-safe pentru mai multe limbaje. Funcționează ca o suită cuprinzătoare pentru validarea configurației, pipeline-uri de serializare a datelor și generarea de cod pentru a asigura consistența datelor pe diferite platforme. Sistemul dispune de un generator de cod multi-limbaj care produce clase de date puternic tipizate din scheme, eliminând nevoia de reflexie. Include un manager de localizare pentru exportul textului tradus și al activelor cu patching specific localității, precum și un pipeline de serializare care transformă fișierele sursă structurate în binare pentru transfer eficient în rețea. Toolchain-ul oferă un motor de validare a configurației pentru a efectua verificări de integritate referențială și verificare a căilor resurselor. De asemenea, suportă modelarea complexă a datelor printr-un sistem de tipuri orientat pe obiecte care permite moștenirea datelor și structuri imbricate în fișierele de configurare.

    Checks referential integrity and resource paths in configuration files to prevent runtime crashes and data errors.

    C#cocos2d-xconfigcsv
    Vezi pe GitHub↗4,458
  • unionai-oss/panderaAvatar unionai-oss

    unionai-oss/pandera

    4,382Vezi pe GitHub↗

    Pandera is a data pipeline validation framework and statistical type validation tool. It functions as a library for defining and enforcing schemas on datasets to ensure data quality and consistency, specifically providing validation capabilities for Pandas dataframes. The project includes a schema inference tool that automates setup by analyzing existing dataset samples to generate validation schemas. It also serves as a synthetic data generator, creating artificial datasets based on predefined schemas to verify data-producing functions. The framework covers data engineering quality assuranc

    Data validation through declarative schemas.

    Pythonassertionsdata-assertionsdata-check
    Vezi pe GitHub↗4,382
  • pyeve/cerberusAvatar pyeve

    pyeve/cerberus

    3,284Vezi pe GitHub↗

    Lightweight, extensible data validation library for Python

    Data validation through schemas.

    Pythondata-validationpython
    Vezi pe GitHub↗3,284
  • danielbeach/data-engineering-practiceAvatar danielbeach

    danielbeach/data-engineering-practice

    2,726Vezi pe GitHub↗

    Data engineering practice repository providing tutorials, distributed processing engines, and Python data pipeline automation scripts. The system encompasses automated data validation, distributed compute aggregation, embedded columnar querying, lazy evaluation planning, partitioned storage export, and cloud storage retrieval. The capability surface covers cloud integration and storage, data engineering and pipelines, data processing and analytics, data quality and testing, database and storage, file management, and monitoring and observability.

    Validates incoming datasets against automated schema rules and structural constraints prior to pipeline execution.

    Python
    Vezi pe GitHub↗2,726
  • seldonio/alibi-detectAvatar SeldonIO

    SeldonIO/alibi-detect

    2,523Vezi pe GitHub↗

    Algorithms for outlier, adversarial and drift detection

    Outlier, adversarial, and drift detection.

    Jupyter Notebook
    Vezi pe GitHub↗2,523
  • frappe/crmAvatar frappe

    frappe/crm

    2,363Vezi pe GitHub↗

    Acest proiect este o platformă open-source de gestionare a relațiilor cu clienții (CRM) care funcționează ca un framework de dezvoltare de aplicații low-code. Oferă o interfață unificată pentru urmărirea pipeline-urilor de vânzări, gestionarea interacțiunilor cu clienții și automatizarea rutării lead-urilor. Platforma este construită pentru a servi drept instrument de automatizare a proceselor de business, permițând utilizatorilor să definească structuri de date și fluxuri de lucru personalizate pentru a eficientiza sarcinile operaționale. Sistemul se distinge prin arhitectura sa bazată pe metadate, care permite generarea dinamică de formulare și modelarea documentelor relaționale. Prin utilizarea injecției de scripturi pe partea de server și a scriptării personalizate a formularelor, utilizatorii pot impune logică de business complexă și reguli de validare a datelor direct în aplicație. De asemenea, se integrează cu sisteme de planificare a resurselor întreprinderii (ERP) și sisteme de comunicare externe, centralizând înregistrările financiare și canalele de mesagerie într-un singur tablou de bord de gestionare. Platforma susține o gamă largă de capabilități operaționale, inclusiv alocarea automatizată a lead-urilor, impunerea acordurilor privind nivelul serviciilor (SLA) și comunicarea multi-canal. Utilizatorii pot organiza datele prin vizualizări și perspective personalizate, în timp ce framework-ul subiacent facilitează colaborarea în echipă prin note partajate și istoricul sarcinilor. Sistemul este conceput pentru implementare în medii containerizate, asigurând performanțe consistente pe infrastructură privată sau bazată pe cloud.

    Applies conditional logic and mandatory field checks to ensure data integrity before saving records.

    Vuecrmcrm-connectionscrm-platform
    Vezi pe GitHub↗2,363
  • nathanepstein/doraAvatar nathanepstein

    nathanepstein/dora

    648Vezi pe GitHub↗

    Tools for exploratory data analysis in Python

    Automate EDA, preprocessing, and feature engineering.

    Python
    Vezi pe GitHub↗648
  1. Home
  2. Part of an Awesome List
  3. Databases & Data
  4. Data Quality and Validation

Explorează sub-etichetele

  • Feature Drift DetectorsValidates feature data against baseline profiles to detect drift or quality regressions before they affect models. **Distinct from Data Quality and Validation:** Distinct from Data Quality and Validation: focuses on drift detection and validation of ML features specifically, not general data integrity or schema compliance.
  • Game Data ValidationChecking referential integrity and resource paths in game configuration files to prevent crashes. **Distinct from Data Quality and Validation:** Specific to game configuration files and resource paths rather than general data quality frameworks.