Tools for exploratory data analysis in Python
nathanepstein/dora 的主要功能包括:Data Analysis Tools, Data Quality and Validation。
nathanepstein/dora 的开源替代品包括: danielbeach/data-engineering-practice — Data engineering practice repository providing tutorials, distributed processing engines, and Python data pipeline… frappe/crm — This project is an open-source customer relationship management platform that functions as a low-code application… apache/seatunnel — SeaTunnel is a distributed data integration engine designed to synchronize structured and unstructured data across… blaze/blaze — NumPy and Pandas interface to Big Data. capitalone/datacompy — Pandas, Polars, Spark, and Snowpark DataFrame comparison for humans and more! apache/incubator-superset — This project is a business intelligence suite and SQL data visualization platform used for data analysis, reporting,…
Data engineering practice repository providing tutorials, distributed processing engines, and Python data pipeline automation scripts. The system encompasses automated data validation, distributed compute aggregation, embedded columnar querying, lazy evaluation planning, partitioned storage export, and cloud storage retrieval. The capability surface covers cloud integration and storage, data engineering and pipelines, data processing and analytics, data quality and testing, database and storage, file management, and monitoring and observability.
This project is an open-source customer relationship management platform that functions as a low-code application development framework. It provides a unified interface for tracking sales pipelines, managing customer interactions, and automating lead routing. The platform is built to serve as a business process automation tool, allowing users to define custom data structures and workflows to streamline operational tasks. The system distinguishes itself through its metadata-driven architecture, which enables dynamic form generation and relational document modeling. By utilizing server-side scr
SeaTunnel is a distributed data integration engine designed to synchronize structured and unstructured data across diverse sources and sinks. It functions as a multi-engine execution framework that can run data integration tasks across different distributed computing backends to optimize workload performance. The project is distinguished by a visual data pipeline designer for configuring workflows without manual code and a specialized change data capture tool for streaming incremental database updates. It also includes an enrichment pipeline that integrates large language models and embedding