awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
ydataai avatar

ydataai/pandas-profiling

0
View on GitHub↗
13,610 स्टार्स·1,792 फोर्क्स·Python·MIT·13 व्यूज़docs.sdk.ydata.ai↗

Pandas Profiling

This project is an exploratory data analysis framework and profiling tool designed to generate comprehensive statistical reports from Pandas and Spark DataFrames. It functions as a data quality profiler that identifies missing values, duplicates, and high correlations within tabular datasets.

The tool distinguishes itself through specialized capabilities for time-series analysis, extracting temporal statistics, seasonality, and auto-correlation plots. It also includes a dataset comparison utility to identify structural or content changes between different versions of a dataset.

The analysis surface covers automated data profiling, quality assessment, and unstructured content analysis. Results can be exported as standalone HTML files, JSON formats, or interactive notebook widgets.

A command-line interface is available to process CSV files and generate reports using configuration files.

Features

  • Comprehensive Profiling Reports - Produces comprehensive exploratory data analysis reports detailing dataset statistics and quality.
  • Automated Exploratory Analysis - Provides a framework that automatically generates statistical summaries and visual insights from tabular datasets.
  • Data Profiling - Produces standardized HTML or JSON reports of dataset characteristics with minimal code.
  • Data Quality - Detects missing values, duplicates, and high correlations to ensure data integrity before processing.
  • Profiling Reports - Generates detailed exploratory data analysis reports and descriptive statistics for Pandas and Spark DataFrames.
  • Data Quality Monitors - Identifies problematic data patterns including missing values, high correlation, and duplicates.
  • Agnostic Interfaces - Implements a unified interface that allows the same analysis logic to run on both Pandas and Spark dataframes.
  • Data Quality Profilers - Identifies missing values, duplicates, and high correlations within large tabular datasets.
  • HTML Analysis Reports - Generates portable standalone HTML reports containing embedded visualizations and statistical summaries.
  • Profiling Tools - A specialized profiling tool for extracting temporal statistical information, including auto-correlation, seasonality, and ACF/PACF plots.
  • Statistical Analysis Libraries - Provides a comprehensive set of tools for calculating descriptive statistics and correlations across datasets.
  • Tabular Data Type Inference - Automatically detects column data types in tabular data to select the most appropriate statistical analysis methods.
  • Dataset Comparators - Generates comparative analyses between multiple versions of a dataset to identify structural or content changes.
  • Time Series Analysis - Applies specialized temporal logic to identify seasonality and autocorrelation in time-indexed columns.
  • Time Series Analysis Tools - Extracts temporal statistics, seasonality, and auto-correlation plots from time-dependent data.
  • Data Exploration - Generates HTML profiling reports for dataframes.

स्टार हिस्ट्री

ydataai/pandas-profiling के लिए स्टार हिस्ट्री चार्टydataai/pandas-profiling के लिए स्टार हिस्ट्री चार्ट

AI सर्च

और अधिक बेहतरीन रिपॉजिटरी खोजें

अपनी ज़रूरत को सरल भाषा में बताएं — AI हजारों क्यूरेटेड ओपन-सोर्स प्रोजेक्ट्स को प्रासंगिकता के आधार पर रैंक करता है।

Start searching with AI

Pandas Profiling के ओपन-सोर्स विकल्प

समान ओपन-सोर्स प्रोजेक्ट्स, जो Pandas Profiling के साथ साझा की गई सुविधाओं के आधार पर रैंक किए गए हैं।
  • data-centric-ai-community/fg-data-profilingData-Centric-AI-Community का अवतार

    Data-Centric-AI-Community/fg-data-profiling

    13,609GitHub पर देखें↗

    This project is a data profiling and exploratory data analysis tool designed to generate automated quality reports for Pandas and Spark dataframes. It serves as a system for computing descriptive statistics, identifying correlations, and analyzing univariate and multivariate data patterns. The tool provides specialized capabilities for comparing different versions of datasets to identify changes in data quality and distributions. It includes a dedicated profiler for time-dependent data to extract statistical information such as seasonality and auto-correlation. The software covers a broad an

    Python
    GitHub पर देखें↗13,609
  • pandas-profiling/pandas-profilingpandas-profiling का अवतार

    pandas-profiling/pandas-profiling

    13,609GitHub पर देखें↗

    This project is an exploratory data analysis library and profiling tool for Pandas and Spark DataFrames. It automates the initial investigation of datasets by generating comprehensive descriptive analysis reports, statistical summaries, and data quality warnings. The system functions as a data quality profiler to detect missing values, duplicate rows, and type inconsistencies. It includes a dataset comparison tool for identifying structural and content shifts between different versions of the same data, as well as specialized tools for time-series analysis to calculate auto-correlation and se

    Python
    GitHub पर देखें↗13,609
  • ydataai/ydata-profilingydataai का अवतार

    ydataai/ydata-profiling

    13,388GitHub पर देखें↗

    Ydata-profiling is an automated exploratory data analysis framework designed to generate comprehensive statistical reports and visual summaries from dataframes. It functions as a diagnostic tool for assessing data quality, identifying missing values, duplicates, and outliers, while providing a scalable engine for profiling massive datasets across distributed enterprise environments. The project distinguishes itself through its ability to handle large-scale data through distributed task orchestration and lazy stream processing, which minimizes memory overhead during complex computations. It in

    Pythonbig-data-analyticsdata-analysisdata-exploration
    GitHub पर देखें↗13,388
  • evidentlyai/evidentlyevidentlyai का अवतार

    evidentlyai/evidently

    7,137GitHub पर देखें↗

    Evidently is an AI observability platform and evaluation framework designed to quantify the performance of machine learning models and large language models. It functions as a monitoring tool for detecting data drift and quality degradation in tabular datasets, while providing a specialized analyzer for the faithfulness and correctness of retrieval augmented generation systems. The project distinguishes itself through an evaluation framework that utilizes judge models and custom rubrics to score language model outputs. It includes tools for iterative prompt optimization and the generation of

    Jupyter Notebookdata-driftdata-qualitydata-science
    GitHub पर देखें↗7,137
Pandas Profiling के सभी 30 विकल्प देखें→

अक्सर पूछे जाने वाले प्रश्न

ydataai/pandas-profiling क्या करता है?

This project is an exploratory data analysis framework and profiling tool designed to generate comprehensive statistical reports from Pandas and Spark DataFrames. It functions as a data quality profiler that identifies missing values, duplicates, and high correlations within tabular datasets.

ydataai/pandas-profiling की मुख्य विशेषताएं क्या हैं?

ydataai/pandas-profiling की मुख्य विशेषताएं हैं: Comprehensive Profiling Reports, Automated Exploratory Analysis, Data Profiling, Data Quality, Profiling Reports, Data Quality Monitors, Agnostic Interfaces, Data Quality Profilers।

ydataai/pandas-profiling के कुछ ओपन-सोर्स विकल्प क्या हैं?

ydataai/pandas-profiling के ओपन-सोर्स विकल्पों में शामिल हैं: data-centric-ai-community/fg-data-profiling — This project is a data profiling and exploratory data analysis tool designed to generate automated quality reports for… pandas-profiling/pandas-profiling — This project is an exploratory data analysis library and profiling tool for Pandas and Spark DataFrames. It automates… ydataai/ydata-profiling — Ydata-profiling is an automated exploratory data analysis framework designed to generate comprehensive statistical… evidentlyai/evidently — Evidently is an AI observability platform and evaluation framework designed to quantify the performance of machine… hosseinmoein/dataframe — DataFrame is a C++ tabular data library and manipulation engine designed for managing heterogeneous data in contiguous… data-centric-ai-community/ydata-profiling — This library provides a diagnostic toolkit for automated data profiling and exploratory analysis. It generates…