How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
Data engineering practice repository providing tutorials, distributed processing engines, and Python data pipeline automation scripts. The system encompasses automated data validation, distributed compute aggregation, embedded columnar querying, lazy evaluation planning, partitioned storage export, and cloud storage retrieval. The capability surface covers cloud integration and storage, data engineering and pipelines, data processing and analytics, data quality and testing, database and storage, file management, and monitoring and observability.
This project is an open-source customer relationship management platform that functions as a low-code application development framework. It provides a unified interface for tracking sales pipelines, managing customer interactions, and automating lead routing. The platform is built to serve as a business process automation tool, allowing users to define custom data structures and workflows to streamline operational tasks. The system distinguishes itself through its metadata-driven architecture, which enables dynamic form generation and relational document modeling. By utilizing server-side scr
SeaTunnel is a distributed data integration engine designed to synchronize structured and unstructured data across diverse sources and sinks. It functions as a multi-engine execution framework that can run data integration tasks across different distributed computing backends to optimize workload performance. The project is distinguished by a visual data pipeline designer for configuring workflows without manual code and a specialized change data capture tool for streaming incremental database updates. It also includes an enrichment pipeline that integrates large language models and embedding
Tools for exploratory data analysis in Python
The main features of nathanepstein/dora are: Data Analysis Tools, Data Quality and Validation.
Open-source alternatives to nathanepstein/dora include: danielbeach/data-engineering-practice — Data engineering practice repository providing tutorials, distributed processing engines, and Python data pipeline… frappe/crm — This project is an open-source customer relationship management platform that functions as a low-code application… apache/seatunnel — SeaTunnel is a distributed data integration engine designed to synchronize structured and unstructured data across… blaze/blaze — NumPy and Pandas interface to Big Data. capitalone/datacompy — Pandas, Polars, Spark, and Snowpark DataFrame comparison for humans and more! apache/incubator-superset — This project is a business intelligence suite and SQL data visualization platform used for data analysis, reporting,…