awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to emqx/kuiper

Projects sharing features with Kuiper

30 open-source projects similar to emqx/kuiper, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • apache/flinkapache avatar

    apache/flink

    26,086View on GitHub↗

    Apache Flink is a distributed processing engine designed for both high-throughput, low-latency data streams and finite batch workloads. It functions as a stateful stream processor and a SQL stream processing engine, providing a unified runtime to execute relational queries and event-based transformations. The system is distinguished by its ability to manage persistent operator state to ensure exactly-once processing guarantees and consistency during failures. It features specialized capabilities for complex event processing to detect temporal patterns and handles out-of-order events using eve

    Java
    View on GitHub↗26,086
  • aklivity/zillaaklivity avatar

    aklivity/zilla

    690View on GitHub↗

    🦎 A multi-protocol edge & service proxy. Seamlessly interface web apps, IoT clients, & microservices to Apache Kafka® via declaratively defined, stateless APIs.

    Java
    View on GitHub↗690
  • cocoindex-io/cocoindexcocoindex-io avatar

    cocoindex-io/cocoindex

    6,117View on GitHub↗

    Cocoindex is an incremental data processing engine that builds and maintains live indexes for AI agents, with a core focus on codebase indexing and knowledge graph extraction. The engine uses a function-graph execution model where user-defined Python functions are composed into a directed acyclic graph, and it processes data incrementally so only changed source records or code paths are re-computed, avoiding full recomputation at any scale. It supports automatic schema inference from transformation pipeline type annotations and provides full data lineage tracing, tagging every output record wi

    Rustagentic-data-frameworkaiai-agents
    View on GitHub↗6,117

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • apache/samzaapache avatar

    apache/samza

    842View on GitHub↗

    Mirror of Apache Samza

    Java
    View on GitHub↗842
  • hstreamdb/hstreamhstreamdb avatar

    hstreamdb/hstream

    722View on GitHub↗

    HStreamDB is an open-source, cloud-native streaming database for IoT and beyond. Modernize your data stack for real-time applications.

    Haskell
    View on GitHub↗722
  • pathwaycom/pathwaypathwaycom avatar

    pathwaycom/pathway

    62,959View on GitHub↗

    Pathway is a high-performance data processing framework designed for building unified batch and streaming pipelines. It functions as an orchestrator for complex data transformations, utilizing a differential dataflow engine to process updates incrementally. By treating static datasets and continuous event streams with identical logic, the platform ensures exactly-once processing semantics and consistent results across diverse data sources. The framework distinguishes itself through its specialized support for real-time artificial intelligence and retrieval-augmented generation. It features in

    Pythonbatch-processingdata-analyticsdata-pipelines
    View on GitHub↗62,959
  • apache/sparkapache avatar

    apache/spark

    43,467View on GitHub↗

    Apache Spark is a unified distributed data processing engine designed for large-scale data analysis and computation graphs. It functions as a distributed machine learning framework, a graph processing system, a real-time stream processor, and a SQL analytics engine. The system enables the execution of distributed SQL querying, large-scale graph analysis, and real-time stream analytics across clusters of machines. It also provides a scalable environment for implementing machine learning algorithms and predictive model development on massive datasets. The engine incorporates relational query e

    Scalabig-datajavajdbc
    View on GitHub↗43,467
  • bytewax/bytewaxbytewax avatar

    bytewax/bytewax

    2,022View on GitHub↗

    Python Stream Processing

    Python
    View on GitHub↗2,022
  • risingwavelabs/risingwaverisingwavelabs avatar

    risingwavelabs/risingwave

    9,093View on GitHub↗

    RisingWave is a cloud-native streaming database and real-time analytics engine that uses standard SQL to process continuous data streams. It functions as a streaming data lakehouse, combining the capabilities of a streaming SQL database with a platform that integrates streaming ingestion with open table formats. The system is distinguished by its use of the PostgreSQL wire protocol, allowing it to integrate with existing SQL tools and drivers. It employs a decoupled compute and storage architecture, persisting streaming state and materialized views in cloud object storage to enable independen

    Rustapache-icebergdata-engineeringdatabase
    View on GitHub↗9,093
  • arroyosystems/arroyoArroyoSystems avatar

    ArroyoSystems/arroyo

    4,819View on GitHub↗

    Arroyo is a high-performance stream processing platform built in Rust. It executes continuous SQL queries on streaming data with event-time semantics, enabling accurate windowed aggregations, joins, and stateful computations on unbounded event streams. The platform uses native Rust execution for high throughput and low latency, with periodic checkpointing for exactly-once fault tolerance and horizontal scaling across distributed workers. The system integrates deeply with Kafka for reading and writing topics with exactly-once delivery and supports change data capture (CDC) from MySQL and Postg

    Rustdatadata-stream-processingdev-tools
    View on GitHub↗4,819
  • beava-dev/beavabeava-dev avatar

    beava-dev/beava

    134View on GitHub↗

    Real-time decision features without streaming infra. Turn live events into product reflexes — no Kafka, no Flink, no feature store.

    Rust
    View on GitHub↗134
  • cantino/huginncantino avatar

    cantino/huginn

    49,487View on GitHub↗

    Huginn is an open-source automation platform that functions as an event-driven task automator and webhook integration engine. It enables the creation of agents that monitor web data and automate tasks across various web services, operating as a self-hosted web scraper and JavaScript workflow orchestrator. The system uses a directed graph of event flows to route and transform data between external APIs. It differentiates itself by allowing custom JavaScript execution within workflows to modify data payloads and by integrating human-in-the-loop automation to insert manual judgment or data entry

    Ruby
    View on GitHub↗49,487
  • caskdata/tigoncaskdata avatar

    caskdata/tigon

    284View on GitHub↗

    High Throughput Real-time Stream Processing Framework

    C++
    View on GitHub↗284
  • edwardcapriolo/teknek-coreedwardcapriolo avatar

    edwardcapriolo/teknek-core

    10View on GitHub↗

    The core libraries of the teknek stream processing platform

    Java
    View on GitHub↗10
  • erlio/vernemqerlio avatar

    erlio/vernemq

    3,592View on GitHub↗

    A distributed MQTT message broker based on Erlang/OTP. Built for high quality & Industrial use cases. The VerneMQ mission is active & the project maintained. Thank you for your support!

    Erlang
    View on GitHub↗3,592
  • faust-streaming/faustfaust-streaming avatar

    faust-streaming/faust

    1,874View on GitHub↗

    Python Stream Processing. A Faust fork

    Python
    View on GitHub↗1,874
  • gearpump/gearpumpgearpump avatar

    gearpump/gearpump

    757View on GitHub↗

    Lightweight real-time big data streaming engine over Akka

    Scala
    View on GitHub↗757
  • google/tensorstoregoogle avatar

    google/tensorstore

    1,522View on GitHub↗

    Library for reading and writing large multi-dimensional arrays.

    C++
    View on GitHub↗1,522
  • hailstorm-hs/hailstormhailstorm-hs avatar

    hailstorm-hs/hailstorm

    93View on GitHub↗

    Haskell distributed stream processing with exactly-once semantics

    Haskell
    View on GitHub↗93
  • hazelcast/hazelcast-jethazelcast avatar

    hazelcast/hazelcast-jet

    1,110View on GitHub↗

    Distributed Stream and Batch Processing

    Java
    View on GitHub↗1,110
  • iotsharp/iotsharpIoTSharp avatar

    IoTSharp/IoTSharp

    1,348View on GitHub↗

    IoTSharp is an open-source IoT platform for data collection, processing, visualization, and device management.

    C#
    View on GitHub↗1,348
  • kuzzleio/kuzzlekuzzleio avatar

    kuzzleio/kuzzle

    1,643View on GitHub↗

    Open-source Back-end, self-hostable & ready to use - Real-time, storage, advanced search - Web, Apps, Mobile, IoT -

    JavaScriptapi-serverbackendelasticsearch
    View on GitHub↗1,643
  • lsds/lightsaberlsds avatar

    lsds/LightSaber

    74View on GitHub↗

    Multi-core Window-Based Stream Processing Engine

    C++
    View on GitHub↗74
  • lsds/saberlsds avatar

    lsds/Saber

    44View on GitHub↗

    Window-Based Hybrid CPU/GPU Stream Processing Engine

    Java
    View on GitHub↗44
  • maki-nage/makinagemaki-nage avatar

    maki-nage/makinage

    42View on GitHub↗

    Stream Processing Made Easy

    Python
    View on GitHub↗42
  • mathcoll/t6mathcoll avatar

    mathcoll/t6

    45View on GitHub↗

    t6 is a "Data-first" IoT platform to connect physical Objects with time-series DB and perform Data Analysis.

    JavaScript
    View on GitHub↗45
  • microsoft/trillmicrosoft avatar

    microsoft/Trill

    1,266View on GitHub↗

    Trill is a single-node query processor for temporal or streaming data.

    C#streaming-datatemporal-data
    View on GitHub↗1,266
  • mosaicml/streamingmosaicml avatar

    mosaicml/streaming

    1,521View on GitHub↗

    A Data Streaming Library for Efficient Neural Network Training

    Python
    View on GitHub↗1,521
  • nanomq/nanomqnanomq avatar

    nanomq/nanomq

    2,534View on GitHub↗

    An ultra-lightweight and blazing-fast MQTT Messaging Broker/Bus for IoT Edge & SDV

    Cactor-modelaioasynchronous
    View on GitHub↗2,534
  • nebulastream/nebulastreamnebulastream avatar

    nebulastream/nebulastream

    86View on GitHub↗

    Data management for the sensor-edge-cloud continuum

    C++
    View on GitHub↗86