awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com

Large dataset optimization

Ranking updated Sep 12, 2026

For large dataset optimizations, the first results are druid-io/druid (Druid is a distributed columnar analytical database that provides real-time streaming ingestion and low-latency querying for massive datasets, matching the core requirements for large-scale data processing and storage), clickhouse/clickhouse (ClickHouse is a distributed columnar analytical database built for high-performance, large-scale data processing, real-time query execution, and high-throughput ingestion) and taosdata/tdengine (TDengine is a distributed time-series database and analytics engine engineered for high-speed ingestion and querying of large-scale timestamped data, fitting the need for efficient large-scale data processing and storage systems). citusdata/citus and apache/hadoop round out the shortlist. Compare the match explanations and check the project documentation against your requirements.

Compare open-source tools and libraries for large dataset optimizations on GitHub, categorized to help address performance bottlenecks.

Large dataset optimization

Find the best repos with AI.We'll search the best matching repositories with AI.
  • druid-io/druiddruid-io avatar

    druid-io/druid

    14,020View on GitHub↗

    Druid is a distributed columnar store and online analytical processing database designed for real-time analytics. It functions as a SQL analytics platform and a streaming data ingestion engine, allowing for the analysis of large datasets with low latency to support interactive dashboards and high-concurrency operational workloads. The system integrates a streaming data ingestion engine that loads information via batch or streaming processes to enable immediate analysis of arriving data. It provides high-performance analytical processing to execute slice-and-dice queries on massive data volume

    Druid is a distributed columnar analytical database that provides real-time streaming ingestion and low-latency querying for massive datasets, matching the core requirements for large-scale data processing and storage.

    JavaColumnar DatabasesColumnar Storage Engines
    View on GitHub↗14,020
  • clickhouse/clickhouseClickHouse avatar

    ClickHouse/ClickHouse

    48,229View on GitHub↗

    ClickHouse is a high-performance, columnar analytical database designed for real-time query execution and large-scale data aggregation. It functions as a distributed data warehouse capable of processing petabytes of information, while also providing an embedded engine that integrates directly into applications for native query capabilities without external dependencies. The system is built to handle high-throughput ingestion and complex analytical workloads, delivering millisecond-level latency for interactive dashboards and operational monitoring. The platform distinguishes itself through ad

    ClickHouse is a distributed columnar analytical database built for high-performance, large-scale data processing, real-time query execution, and high-throughput ingestion.

    C++Columnar DatabasesColumnar Storage EnginesQuery Optimization Engines
    View on GitHub↗48,229
  • taosdata/tdenginetaosdata avatar

    taosdata/TDengine

    24,734View on GitHub↗

    TDengine is a distributed time-series database designed for the high-speed ingestion, compression, and retrieval of timestamped metrics and sensor data. It functions as a SQL-compatible analytics engine, allowing users to perform complex operations on massive volumes of time-ordered information using standard relational syntax. The platform is built to serve as a backend foundation for industrial IoT environments, managing real-time data streams and device metadata through a cluster-based architecture. The system distinguishes itself through a distributed sharding architecture that uses consi

    TDengine is a distributed time-series database and analytics engine engineered for high-speed ingestion and querying of large-scale timestamped data, fitting the need for efficient large-scale data processing and storage systems.

    CColumnar Storage EnginesData Compression Algorithms
    View on GitHub↗24,734
  • citusdata/cituscitusdata avatar

    citusdata/citus

    12,562View on GitHub↗

    Citus is a PostgreSQL extension that transforms a standard database into a distributed system. It functions as a sharding framework and distributed SQL engine, enabling horizontal scaling by partitioning tables across a cluster of nodes. By utilizing a coordinator-worker topology, the system manages metadata and routes queries to the appropriate nodes, allowing for parallel execution of complex operations across distributed data shards. The platform distinguishes itself through its specialized support for multi-tenant architectures and real-time analytical processing. It enables tenant-based

    Citus is a distributed PostgreSQL extension that transforms a standard database into a horizontally scalable system capable of distributed computing, query optimization, and handling large-scale datasets.

    CColumnar Storage EnginesData Compression AlgorithmsQuery Optimizers
    View on GitHub↗12,562
  • apache/hadoopapache avatar

    apache/hadoop

    15,567View on GitHub↗

    Hadoop is a big data infrastructure suite and distributed data processing framework designed to store and process massive datasets across clusters of computers. It consists of a distributed storage system for managing large files across multiple nodes and a parallel computing engine for processing data across a distributed cluster. The framework implements a distributed file system to ensure fault tolerance and high throughput, paired with a programming model that processes large datasets in parallel. It manages the underlying hardware and software environment required for distributed big dat

    Hadoop is a foundational distributed data processing framework designed to store and analyze large-scale datasets across clusters, aligning squarely with the required capabilities.

    JavaDistributed Computing
    View on GitHub↗15,567
  • oceanbase/oceanbaseoceanbase avatar

    oceanbase/oceanbase

    9,980View on GitHub↗

    OceanBase is a distributed SQL database designed for high availability and strong consistency across multiple nodes and regions. It functions as a hybrid transactional and analytical processing engine, allowing real-time analytics and transactions to execute on a single data copy. The system also serves as a vector database engine for indexing and querying vector data to power semantic search and recommendation systems. The platform features native compatibility layers for MySQL and Oracle, enabling the migration of legacy workloads without rewriting SQL code. It utilizes a Paxos-based distri

    OceanBase is a distributed SQL database and HTAP engine that supports large-scale data processing with columnar storage, distributed computing, and data compression.

    C++Columnar Storage EnginesData Compression Algorithms
    View on GitHub↗9,980
  • apache/arrowapache avatar

    apache/arrow

    16,529View on GitHub↗

    Arrow is a cross-language development platform for in-memory data. It provides a standardized, language-independent columnar memory format designed to accelerate analytical operations and improve memory efficiency on modern computing hardware. By utilizing a schema-driven approach, the framework enables the efficient organization of both flat and nested data structures. The project functions as an analytical data processing engine that facilitates high-performance computation directly on memory-resident datasets. It distinguishes itself through a zero-copy architecture, which allows multiple

    Apache Arrow provides a standardized columnar memory format and zero-copy architecture designed for high-performance in-memory analytics across multiple languages, fitting the large-scale data processing framework category well.

    C++Columnar Data Processors
    View on GitHub↗16,529
  • trinodb/trinotrinodb avatar

    trinodb/trino

    12,952View on GitHub↗

    Trino is a distributed SQL query engine designed for large-scale data analytics. It functions as a data federation platform, providing a unified interface that allows users to execute complex analytical queries across multiple heterogeneous data sources simultaneously without requiring data movement or transformation. The engine utilizes a massively parallel processing architecture to scale compute resources across clusters for high-speed data retrieval. It distinguishes itself through a cost-based query optimizer that analyzes metadata to determine efficient execution plans, alongside dynami

    Trino is a distributed SQL query engine for large-scale data analytics that fits the category through its massively parallel processing architecture, distributed computing capabilities, and cost-based query optimizer.

    JavaCost-Based Optimizers
    View on GitHub↗12,952
  • finos/perspectivefinos avatar

    finos/perspective

    10,967View on GitHub↗

    Perspective is a columnar data analytics library and streaming data visualization engine. It provides an interactive data grid component and notebook analytics widgets designed for processing high-volume data and rendering interactive charts and grids. The system utilizes a high-performance query engine to enable real-time data analysis and streaming dataset visualization. It supports the creation of customizable dashboards and reports that update automatically as new data arrives without requiring full dataset reloads. The project covers large-scale dataset analytics through a schema-driven

    Perspective is a columnar data analytics library and streaming data visualization engine featuring a high-performance query engine and columnar storage, making it well-suited for high-volume data analysis though it focuses more on interactive visualization and grids than full distributed computing clusters.

    C++Columnar Storage EnginesColumnar Data Processors
    View on GitHub↗10,967
  • apache/hiveapache avatar

    apache/hive

    6,012View on GitHub↗

    Apache Hive is a SQL-on-Hadoop data warehouse that enables querying and managing petabytes of data stored in distributed storage such as HDFS and cloud storage services. It provides a familiar SQL interface for batch analytics and reporting, supported by a core set of components including the HiveServer2 Thrift service for remote query execution, the Hive Metastore Service for central metadata management, the Hive ACID Transaction Engine for concurrent read-write operations, and the Hive LLAP Interactive Engine for low-latency analytical processing. The WebHCat REST API offers an HTTP interfac

    Apache Hive is a distributed data warehouse and query engine designed for petabyte-scale batch analytics and reporting on Hadoop-compatible storage, matching the large-scale data processing and storage framework category.

    JavaColumnar Storage EnginesCost-Based Optimizers
    View on GitHub↗6,012
  • apache/sparkapache avatar

    apache/spark

    43,467View on GitHub↗

    Apache Spark is a unified distributed data processing engine designed for large-scale data analysis and computation graphs. It functions as a distributed machine learning framework, a graph processing system, a real-time stream processor, and a SQL analytics engine. The system enables the execution of distributed SQL querying, large-scale graph analysis, and real-time stream analytics across clusters of machines. It also provides a scalable environment for implementing machine learning algorithms and predictive model development on massive datasets. The engine incorporates relational query e

    Apache Spark is a distributed computing engine that provides in-memory processing, SQL querying, and large-scale data analysis, which squarely fits the requirement for a heavy-duty data processing framework.

    ScalaCost-Based Optimizers
    View on GitHub↗43,467
  • apache/flinkapache avatar

    apache/flink

    26,086View on GitHub↗

    Apache Flink is a distributed processing engine designed for both high-throughput, low-latency data streams and finite batch workloads. It functions as a stateful stream processor and a SQL stream processing engine, providing a unified runtime to execute relational queries and event-based transformations. The system is distinguished by its ability to manage persistent operator state to ensure exactly-once processing guarantees and consistency during failures. It features specialized capabilities for complex event processing to detect temporal patterns and handles out-of-order events using eve

    Apache Flink is a distributed streaming and batch processing engine that provides stateful stream processing, unified relational queries, and large-scale data processing capabilities aligned with this search.

    JavaQuery Optimizers
    View on GitHub↗26,086
  • rapidsai/cudfrapidsai avatar

    rapidsai/cudf

    9,672View on GitHub↗

    cuDF is a GPU-accelerated dataframe library and data processing engine designed for manipulating and analyzing large tabular datasets. It provides a high-level API for executing filtering, joining, and aggregating operations directly on GPU hardware. The project integrates the Apache Arrow memory format to enable zero-copy data transfers and includes a just-in-time compiler for executing custom user-defined functions on the GPU. The library features specialized acceleration for existing workflows by redirecting standard Pandas dataframe calls and Polars query plans to a GPU backend. It also p

    This GPU-accelerated dataframe library provides high-performance data processing and manipulation for large datasets, though it focuses on single-node GPU acceleration rather than full distributed computing across clusters.

    C++Distributed ComputingZero-Copy Data AccessZero-Copy Data Access Libraries
    View on GitHub↗9,672
  • apache/pinotapache avatar

    apache/pinot

    6,098View on GitHub↗

    Pinot is a distributed, columnar analytical database designed for high-concurrency, low-latency query processing. It functions as a real-time OLAP datastore, enabling interactive, user-facing analytics by ingesting and querying massive datasets from both streaming and batch sources. The system architecture relies on a centralized controller for cluster coordination and a distributed segment-based storage model to ensure horizontal scalability. The platform distinguishes itself through a hybrid ingestion pipeline that unifies real-time event streams and historical batch data into a single quer

    Apache Pinot is a distributed columnar analytical database that provides low-latency, high-concurrency query processing for large-scale datasets, making it an ideal fit for the requested data storage and processing framework.

    JavaColumnar Storage EnginesQuery Optimizers
    View on GitHub↗6,098
  • apache/hbaseapache avatar

    apache/hbase

    5,540View on GitHub↗

    HBase is a distributed, wide-column NoSQL store and big data storage engine designed for sparse datasets. It functions as a scalable columnar database built on top of the Hadoop Distributed File System to provide real-time read and write access to massive volumes of structured and unstructured data. The system acts as a cross-language database gateway, offering connectivity through native remote procedure calls, REST, and Thrift interfaces. It distinguishes itself through a master-worker coordination model that enables horizontal scaling and fault tolerance across a cluster. The project cove

    HBase is a distributed, wide-column big data storage engine built for large-scale datasets, making it the right kind of system despite lacking in-memory processing and zero-copy features.

    JavaColumnar Databases
    View on GitHub↗5,540
  • pola-rs/polarspola-rs avatar

    pola-rs/polars

    38,855View on GitHub↗

    Polars is a high-performance columnar data processing library designed for efficient analytical workflows. It functions as a structured data library that organizes information into typed columns, utilizing the Apache Arrow memory format to enable zero-copy data sharing and cache-friendly, vectorized operations. The engine is built to handle large-scale tabular datasets, providing both local and distributed analytical runtimes that scale from single-machine environments to multi-node clusters. The project distinguishes itself through a sophisticated lazy query engine that constructs abstract e

    Polars is a high-performance columnar data processing library featuring zero-copy data sharing via Apache Arrow and distributed runtimes, fitting the large-scale data processing category well despite focusing on tabular dataframes rather than broad distributed storage systems.

    RustColumnar Storage EnginesQuery OptimizersColumnar Data Processors
    View on GitHub↗38,855
  • apache/druidapache avatar

    apache/druid

    14,020View on GitHub↗

    Apache Druid is a real-time analytics database and distributed columnar time-series store designed for sub-second analytical queries. It functions as a data platform featuring a distributed SQL query engine and a real-time data ingestion system for moving historical and streaming data from external sources. The system is distinguished by its ability to provide low-latency analytics under high concurrency to power operational dashboards. It implements a Kerberos-secured environment for user authentication and employs a shared-nothing cluster architecture to enable horizontal scaling. The plat

    Apache Druid is a distributed columnar analytics database that provides fast query performance and real-time data ingestion for large-scale datasets, making it well-suited for high-concurrency analytical workloads despite lacking a few of the requested data processing primitives.

    JavaColumnar Storage EnginesQuery Planning
    View on GitHub↗14,020
  • facebookincubator/veloxfacebookincubator avatar

    facebookincubator/velox

    4,155View on GitHub↗

    Velox is a high-performance C++ query execution engine and columnar data processing library. It serves as a composable framework for implementing analytical query engines, providing a vectorized expression evaluator and a toolkit for data management systems. The project is distinguished by its use of vectorized columnar execution and arena-based memory allocation to process large-scale datasets. It features specialized optimizations such as broadcast join table caching, dynamic filter push-down, and dictionary encoding to reduce memory overhead and accelerate analytical reads. The engine cov

    Velox is a high-performance C++ query execution engine and columnar data processing library that provides vectorized execution and memory management for large-scale datasets, though as a building block for data systems rather than a complete standalone distributed computing framework.

    C++Columnar Data ProcessorsNested Columnar StorageNested Columnar Storage
    View on GitHub↗4,155
  • eto-ai/lanceeto-ai avatar

    eto-ai/lance

    6,671View on GitHub↗

    Lance is a versioned columnar data format and storage engine designed as a multimodal AI lakehouse. It serves as a vector database storage engine and a cloud object store dataset manager, organizing images, video, audio, and embeddings into a unified format optimized for machine learning workflows. The project distinguishes itself by combining a columnar layout for structured data with a specialized blob store for large multimodal tensors. It implements a hybrid search engine that integrates vector similarity search, full-text search, and SQL analytics on a single dataset, supported by a stor

    Lance is a columnar data format and storage engine built for large-scale multimodal datasets, featuring in-memory caching and vectorized queries that fit well for high-performance data processing tasks.

    RustColumnar Storage EnginesNested Columnar Storage
    View on GitHub↗6,671
  • vaexio/vaexvaexio avatar

    vaexio/vaex

    8,506View on GitHub↗

    Vaex is a high-performance Apache Arrow DataFrame library and out-of-core data processing engine designed to handle billion-row tabular datasets in Python. It functions as a lazy evaluation framework that defers computations and transformations until results are required, enabling the processing of datasets that exceed available system RAM by mapping files directly from disk. The project distinguishes itself as a tool for big data visualization and exploration, specifically integrated for use within interactive notebooks. It provides specialized capabilities for machine learning feature engin

    Vaex is a high-performance out-of-core DataFrame library and data processing engine that uses memory-mapped files and lazy evaluation to analyze massive tabular datasets efficiently in Python, making it a strong fit despite lacking a fully distributed computing architecture.

    PythonColumnar Tabular Storage
    View on GitHub↗8,506
  • lancedb/lancedblancedb avatar

    lancedb/lancedb

    9,031View on GitHub↗

    LanceDB is a vector database and columnar data store designed to function as a versioned dataset manager and vector search engine. It serves as a high-performance backend for indexing and retrieving high-dimensional embeddings, providing the foundation for machine learning data pipelines. The system distinguishes itself through a combination of cloud-native object storage and immutable version tracking, allowing for data time-travel and reproducible AI experiments. It integrates hybrid search capabilities, merging dense vector similarity with BM25 full-text search and SQL-like scalar filters

    LanceDB is a columnar data store and vector database built for high-performance retrieval and dataset management, making it a strong fit for large-scale data processing despite its primary focus on vector search.

    HTMLColumnar DatabasesColumnar Storage EnginesQuery Planning
    View on GitHub↗9,031
  • eventual-inc/daftEventual-Inc avatar

    Eventual-Inc/Daft

    5,225View on GitHub↗

    Daft is a distributed dataframe library and multimodal data processor designed to handle large-scale structured and unstructured data. It functions as a vectorized execution engine that processes tables alongside images, audio, and video, utilizing a unified schema to manage diverse data types. The project distinguishes itself by combining distributed data engineering with large-scale AI inference. It provides an AI data pipeline for batch-optimizing model prompts and generating high-dimensional text embeddings, while utilizing zero-copy memory sharing to execute custom Python functions witho

    Daft is a distributed dataframe and multimodal data processor designed to handle large-scale datasets efficiently, featuring distributed computing, zero-copy memory sharing, and support for structured and unstructured data pipelines.

    RustDistributed Computing
    View on GitHub↗5,225
  • prestodb/prestoprestodb avatar

    prestodb/presto

    16,711View on GitHub↗

    Presto is a distributed SQL query engine designed for high-performance analytical processing across heterogeneous data sources. It functions as a data federation platform and massively parallel processing engine, allowing users to execute interactive queries against diverse storage systems without requiring data migration. By mapping remote metadata and structures to a unified relational namespace, it enables seamless cross-platform analysis through a standard SQL interface. The engine distinguishes itself through a pluggable connector architecture and a shared-nothing distributed processing

    Presto is a distributed SQL query engine designed for high-performance analytical processing across heterogeneous data sources, providing the distributed computing capability needed for large-scale data analysis despite missing a few specific storage-layer features.

    JavaQuery OptimizersCost-Based OptimizersQuery Planning
    View on GitHub↗16,711
  • dask/daskdask avatar

    dask/dask

    13,746View on GitHub↗

    Dask is a parallel computing framework and distributed task scheduler designed to scale Python data science workflows from single machines to large clusters. It functions as a cluster resource manager that orchestrates computational logic by representing tasks and their dependencies as directed acyclic graphs. This architecture allows the system to automate the distribution of workloads across available hardware while managing complex execution requirements. The project distinguishes itself through a lazy evaluation engine that defers data operations until they are explicitly requested, enabl

    Dask is a parallel computing framework that scales Python data science workflows across clusters, though it lacks dedicated columnar storage formats and built-in zero-copy deserialization capabilities.

    PythonDistributed ComputingQuery Optimizations
    View on GitHub↗13,746
  • greptimeteam/greptimedbGreptimeTeam avatar

    GreptimeTeam/greptimedb

    5,968View on GitHub↗

    GreptimeDB is a distributed, open-source time-series database built for unified observability. It stores and queries metrics, logs, and traces together in a single columnar engine, supporting both SQL and PromQL for analysis. The database is designed as a Kubernetes-native operator with a decoupled compute and storage architecture, enabling horizontal scaling and multi-region deployment. What distinguishes GreptimeDB is its role as a multi-protocol ingestion gateway, accepting data through OpenTelemetry, Prometheus Remote Write, InfluxDB, Loki, Elasticsearch, Kafka, and MQTT protocols without

    GreptimeDB is a distributed time-series database featuring a columnar storage engine and decoupled compute-storage architecture that handles large-scale metrics, logs, and traces efficiently, though it is specifically tailored for time-series data rather than general-purpose big data processing.

    RustColumnar Storage EnginesData Compression Algorithms
    View on GitHub↗5,968
  • ray-project/rayray-project avatar

    ray-project/ray

    42,895View on GitHub↗

    Ray is a distributed computing framework designed to scale Python and Java applications across clusters by abstracting task scheduling and resource management. It functions as a resource-aware execution engine that manages task dependencies, placement, and fault tolerance across networked compute nodes. At its core, the system provides a stateful actor model, allowing developers to define classes that run in dedicated processes to maintain and mutate internal state across remote method calls. The framework distinguishes itself through a robust cross-language interoperability layer, enabling f

    Ray is a distributed computing framework that provides parallel task scheduling and stateful actor execution for scaling workloads across clusters, fitting the required distributed computing domain.

    PythonQuery Optimization Engines
    View on GitHub↗42,895
  • tporadowski/redistporadowski avatar

    tporadowski/redis

    9,987View on GitHub↗

    Redis is a high-performance in-memory key-value store that functions as a distributed cache, message broker, and NoSQL database. It provides sub-millisecond read and write access to data stored in RAM and can operate as a vector database for indexing high-dimensional embeddings. The system supports a wide range of data storage and synchronization primitives, including the management of strings, hashes, lists, sets, and JSON documents. It enables real-time data operations through atomic transactions, hybrid persistence using snapshots and append-only logs, and high-availability configurations

    Redis provides high-performance, in-memory data storage and caching with support for clustering and replication, though it is primarily designed as an in-memory key-value store rather than a distributed analytical data processing framework.

    CDistributed CachesIn-Memory Data StoresKey-Value Stores
    View on GitHub↗9,987
  • apache/cassandraapache avatar

    apache/cassandra

    9,778View on GitHub↗

    Cassandra is a distributed NoSQL database and wide-column store designed for high availability and linear scalability. It functions as a fault-tolerant distributed system that utilizes an LSM-tree storage engine to optimize write throughput and manage massive datasets. The system is a CQL-compliant database, using a structured query language to manage and retrieve tabular data stored across multiple nodes. It organizes information into rows and columns based on a flexible schema and primary keys. The project provides capabilities for horizontal database scaling, distributed data partitioning

    Apache Cassandra is a distributed wide-column NoSQL database that handles massive datasets with linear scalability, making it a strong fit for distributed data storage despite lacking the broader in-memory compute frameworks typical of general data processing engines.

    JavaNoSQL DatabasesWide-Column StoresConsistent Hashing
    View on GitHub↗9,778
  • apache/icebergapache avatar

    apache/iceberg

    8,972View on GitHub↗

    Iceberg is an open table format and big data table manager designed for huge analytic datasets in cloud storage. It provides a specification for tracking large-scale datasets to maintain transactional consistency and structural integrity. The project utilizes a standardized REST catalog interface to manage table metadata, ensuring interoperability between different compute engines. This allows diverse query engines to connect to a single table interface and maintain consistency across different processing frameworks. Its core capabilities include managing large-scale analytic tables, coordin

    Apache Iceberg is an open table format and big data table manager built for handling huge analytic datasets across multiple compute engines, fitting the large-scale data processing and storage framework category despite lacking some built-in processing features.

    JavaBig Data and AnalyticsTable ManagersAtomic Write Coordinators
    View on GitHub↗8,972
  • delta-io/deltadelta-io avatar

    delta-io/delta

    8,596View on GitHub↗

    Delta is a lakehouse table format that brings ACID transactions and data warehouse consistency to large scale data lakes on cloud object storage. It serves as an ACID transaction manager, coordinating atomic commits and serializable isolation for concurrent reads and writes across distributed compute engines. The project provides a multi-engine interoperability layer that uses format translation to allow diverse SQL engines and processing frameworks to read and write the same tables. It functions as a data versioning system, utilizing a transaction log to enable time travel, historical snapsh

    Delta is a storage layer and table format that brings ACID transactions and reliability to large-scale data lakes, fitting the domain of data processing frameworks even though it relies on external distributed engines for execution.

    ScalaLakehouse Storage LayersLakehouse Table FormatsTransaction Management
    View on GitHub↗8,596
  • apache/hudiapache avatar

    apache/hudi

    6,097View on GitHub↗

    Apache Hudi is an open-source table format that brings ACID transactions, incremental processing, and multi-modal indexing to data lakes. It provides atomic commits with snapshot isolation, rollback, and optimistic concurrency control for reliable data lake operations, while supporting upserts, record-level updates, and deletions in large analytical datasets. The project distinguishes itself through a timeline-based architecture that coordinates all write operations, enabling features like time-travel querying, incremental change streaming, and multi-modal query views that include snapshot, i

    Apache Hudi is a data lake table format and storage layer designed to bring ACID transactions and incremental processing to large-scale analytical datasets, aligning well with distributed data processing needs even though it focuses on table management rather than a complete compute framework.

    JavaTransactional Data Lake EnginesACID Transactional CoresBatch Table Ingesters
    View on GitHub↗6,097
  • modin-project/modinmodin-project avatar

    modin-project/modin

    10,389View on GitHub↗

    Modin is a distributed dataframe library and parallel data processing engine designed to handle large datasets that exceed system memory. It functions as a distributed computing framework that parallelizes data manipulation tasks across multiple CPU cores or clusters to increase throughput and avoid memory errors. The project mirrors the Pandas API, allowing for the distribution of data workflows without changing core code logic. It utilizes a pluggable backend interface, which enables users to switch between different distributed execution engines to optimize performance based on available h

    Modin is a distributed dataframe library that parallelizes pandas workflows to process large-scale datasets across multiple cores or clusters, making it a well-suited tool for large-scale data processing despite lacking some standalone storage formats like columnar layouts.

    PythonDistributed Compute FrameworksDistributed Data Processing FrameworksAPI Compatibility Layers
    View on GitHub↗10,389
  • huggingface/datasetshuggingface avatar

    huggingface/datasets

    21,643View on GitHub↗

    Datasets is a library designed for the management, processing, and sharing of large-scale data collections for machine learning workflows. It functions as both a data processing framework and a versioning platform, providing tools to organize, filter, and transform massive datasets while ensuring reproducibility across research and development teams. The library distinguishes itself by enabling the handling of datasets that exceed available system memory. It utilizes memory-mapped file access, disk-based caching, and lazy iterative streaming to maintain performance when working with large-sca

    This library functions as a data processing framework for large-scale machine learning datasets, utilizing memory-mapped storage and lazy streaming to handle data exceeding system memory, though it lacks general-purpose distributed computing features.

    PythonPython Machine Learning LibrariesData Processing FrameworksDataset Versioning Platforms
    View on GitHub↗21,643
Compare the top 10 at a glance
RepositoryStarsLanguageLicenseLast push
druid-io/druid14KJavaApache-2.0Jun 17, 2026
clickhouse/clickhouse48.2KC++Apache-2.0Jun 23, 2026
taosdata/tdengine
24.7K
C
agpl-3.0
Feb 21, 2026
citusdata/citus12.6KCAGPL-3.0Jun 16, 2026
apache/hadoop15.6KJavaApache-2.0Jun 17, 2026
oceanbase/oceanbase10KC++otherFeb 14, 2026
apache/arrow16.5KC++apache-2.0Feb 21, 2026
trinodb/trino13KJavaApache-2.0Jun 23, 2026
finos/perspective11KC++Apache-2.0Jun 5, 2026
apache/hive6KJavaapache-2.0Feb 20, 2026

Related searches

  • Sprite optimization tool
  • Vector indexing database
  • Data pagination library
  • Mathematical optimization engine
  • a dataframe engine for huge data
  • JavaScript bundle optimizer
  • Data partitioning strategies
  • a high performance library for tabular data