For an analytics database for slicing huge tables fast, the strongest matches are apache/doris (Doris is a distributed SQL data warehouse built on), starrocks/starrocks (StarRocks is a distributed SQL OLAP database with a) and druid-io/druid (Druid is a distributed columnar OLAP database purpose-built for). databendlabs/databend and questdb/questdb round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
High-performance open-source database systems optimized for rapid analytical processing and complex large-scale data queries.
Doris is a distributed SQL data warehouse designed for high-performance analytical workloads and real-time data processing. It functions as a unified platform that integrates traditional relational warehousing with lakehouse query capabilities, allowing users to execute analytical operations directly against external data lakes without requiring data migration. The system distinguishes itself through a shared-nothing, massively parallel processing architecture that utilizes vectorized query execution and columnar storage to maintain sub-second latency. It supports dynamic schema evolution, en
Doris is a distributed SQL data warehouse built on a shared-nothing MPP architecture with columnar storage and vectorized query execution, exactly matching your requirement for a scalable OLAP database with fast analytical query capabilities.
StarRocks is a distributed SQL OLAP database engine designed for real-time analytics and high-performance multi-dimensional analysis. It functions as a data lakehouse query engine that enables SQL execution across large datasets and external open table formats without requiring local data imports. The system employs a shared-nothing distributed architecture and utilizes the MySQL protocol to integrate with business intelligence tools. It maintains real-time data consistency through a primary key upsert model and accelerates query response times using vectorized execution and cost-based optimi
StarRocks is a distributed SQL OLAP database with a shared-nothing MPP architecture, vectorized execution, and full SQL support via MySQL protocol, directly matching your need for a columnar database that handles fast analytical queries at scale.
Druid is a distributed columnar store and online analytical processing database designed for real-time analytics. It functions as a SQL analytics platform and a streaming data ingestion engine, allowing for the analysis of large datasets with low latency to support interactive dashboards and high-concurrency operational workloads. The system integrates a streaming data ingestion engine that loads information via batch or streaming processes to enable immediate analysis of arriving data. It provides high-performance analytical processing to execute slice-and-dice queries on massive data volume
Druid is a distributed columnar OLAP database purpose-built for real-time, fast analytical queries at scale, making it a flagship answer for this search.
Databend is a cloud-native data warehouse and OLAP database designed for large-scale analytics. It functions as a SQL-compliant engine and serverless analytics platform that separates compute from storage to allow for independent scaling. The system integrates vector database capabilities, indexing high-dimensional embeddings to enable semantic, hybrid, and full-text searches across massive datasets. It further distinguishes itself through serverless compute management that automatically scales resources based on demand and shuts them down during idle periods. The platform covers a broad set
Databend is a cloud-native OLAP database and data warehouse with SQL compliance, massively parallel processing, and horizontal scalability, directly matching the need for scalable analytical querying.
QuestDB is a high-performance, distributed time-series database designed for the ingestion, storage, and analysis of massive datasets. It functions as a real-time analytics platform that utilizes a columnar storage engine to optimize disk input and output, enabling efficient analytical scans and complex windowing operations on streaming data. The platform distinguishes itself through specialized capabilities for handling asynchronous time-series streams, including advanced join algorithms that align disparate data sets based on precise timestamp lookups. It supports high-volume ingestion thro
QuestDB is a high-performance distributed time-series database with a columnar storage engine, SQL support, SIMD vectorized execution, and horizontal scalability, making it a strong match for an open-source columnar OLAP database that can handle analytical queries at scale.
DuckDB is an embedded, in-process analytical SQL database and OLAP database management system. It functions as a data engine for Parquet and CSV files, allowing users to execute complex SQL queries on large datasets without requiring a separate server process. The system is designed for local analytical processing and embedded data science workflows. It enables the direct querying and analysis of Parquet and CSV files from disk, bypassing the need to load data into a permanent database. The engine provides high-performance analytical SQL execution, including support for window functions and
DuckDB is an open-source embedded columnar OLAP database that delivers fast analytical queries with vectorized execution, MPP, compression, and ANSI SQL, though its in-process design means it lacks the horizontal node scalability of a distributed system.
ClickHouse is a high-performance, columnar analytical database designed for real-time query execution and large-scale data aggregation. It functions as a distributed data warehouse capable of processing petabytes of information, while also providing an embedded engine that integrates directly into applications for native query capabilities without external dependencies. The system is built to handle high-throughput ingestion and complex analytical workloads, delivering millisecond-level latency for interactive dashboards and operational monitoring. The platform distinguishes itself through ad
ClickHouse is an open-source columnar OLAP database that delivers high-performance analytical queries at scale, with support for massively parallel processing, data compression, ANSI SQL, horizontal scalability, and vectorized execution — exactly matching this search.
DuckDB is an in-process analytical database engine designed to run directly within an application process. As a zero-dependency, embedded system, it provides enterprise-grade SQL data processing capabilities without the overhead of managing a dedicated database server. It is built to handle complex analytical and aggregation tasks by storing and retrieving information in columns, allowing for high-performance relational data manipulation. The engine distinguishes itself through a columnar vectorized execution model that maximizes CPU cache efficiency during query operations. It employs adapti
DuckDB is a genuine columnar OLAP database with vectorized execution and SQL support, but it is an embedded engine without distributed MPP or horizontal scalability, so it fits the category but misses some large-scale features.
Apache Druid is a real-time analytics database and distributed columnar time-series store designed for sub-second analytical queries. It functions as a data platform featuring a distributed SQL query engine and a real-time data ingestion system for moving historical and streaming data from external sources. The system is distinguished by its ability to provide low-latency analytics under high concurrency to power operational dashboards. It implements a Kerberos-secured environment for user authentication and employs a shared-nothing cluster architecture to enable horizontal scaling. The plat
Apache Druid is a real-time distributed columnar OLAP database with a SQL query engine and horizontal scaling, fitting your search for a columnar analytics database, though its time-series focus makes it narrower than a general-purpose OLAP system.
OceanBase is a distributed SQL database designed for high availability and strong consistency across multiple nodes and regions. It functions as a hybrid transactional and analytical processing engine, allowing real-time analytics and transactions to execute on a single data copy. The system also serves as a vector database engine for indexing and querying vector data to power semantic search and recommendation systems. The platform features native compatibility layers for MySQL and Oracle, enabling the migration of legacy workloads without rewriting SQL code. It utilizes a Paxos-based distri
OceanBase is a distributed SQL database with columnar storage and HTAP capabilities that supports analytical queries at scale, meeting the key requirements for an OLAP database even though it also handles transactional workloads.
Pinot is a distributed, columnar analytical database designed for high-concurrency, low-latency query processing. It functions as a real-time OLAP datastore, enabling interactive, user-facing analytics by ingesting and querying massive datasets from both streaming and batch sources. The system architecture relies on a centralized controller for cluster coordination and a distributed segment-based storage model to ensure horizontal scalability. The platform distinguishes itself through a hybrid ingestion pipeline that unifies real-time event streams and historical batch data into a single quer
Pinot is a distributed, columnar OLAP database built for real-time, low-latency analytical queries at scale, with columnar storage, horizontal scalability, SQL support, and a segment-based MPP architecture, squarely matching the core requirements of this search.
Apache Druid is a real-time OLAP database and distributed analytics engine. It functions as a columnar time-series database designed for high-performance analytical queries and the real-time ingestion of streaming and batch datasets. The system provides a framework for high-concurrency analytics, allowing multiple simultaneous users to execute SQL and native queries across large-scale data. It supports mixed data ingestion, combining real-time streaming and batch loading into a single system for unified analysis. The platform includes capabilities for distributed cluster management, enabling
Apache Druid is a real-time columnar OLAP database built for high-performance analytical queries on streaming and batch data, with distributed SQL and parallel execution — it fits the core need for a columnar OLAP database at scale, though its description does not explicitly highlight data compression or vectorized query execution.
| Repository | Stars | Language | License | Last push |
|---|---|---|---|---|
| apache/doris | 15.5K | Java | Apache-2.0 | |
| starrocks/starrocks | 11.8K | Java | Apache-2.0 | |
| druid-io/druid | 14K | Java | Apache-2.0 | |
| databendlabs/databend | 9.4K | Rust | NOASSERTION | |
| questdb/questdb | 17.1K | Java | Apache-2.0 | |
| cwida/duckdb | 38.8K | C++ | MIT | |
| clickhouse/clickhouse | 48.2K | C++ | Apache-2.0 | |
| duckdb/duckdb | 38.8K | C++ | MIT | |
| apache/druid | 14K | Java | Apache-2.0 | |
| oceanbase/oceanbase | 10K | C++ | other |