awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
apache avatar

apache/samza

0
View on GitHub↗
842 stars·332 forks·Java·Apache-2.0·34 views

Samza

Mirror of Apache Samza

Features

  • Stream Processing - Provides distributed stream processing with fault tolerance.
  • Streaming Engines - Distributed framework built on Kafka and YARN.

Star history

Star history chart for apache/samzaStar history chart for apache/samza

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Samza

These projects share indexed features with Samza. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • apache/sparkapache avatar

    apache/spark

    43,467View on GitHub↗

    Apache Spark is a unified distributed data processing engine designed for large-scale data analysis and computation graphs. It functions as a distributed machine learning framework, a graph processing system, a real-time stream processor, and a SQL analytics engine. The system enables the execution of distributed SQL querying, large-scale graph analysis, and real-time stream analytics across clusters of machines. It also provides a scalable environment for implementing machine learning algorithms and predictive model development on massive datasets. The engine incorporates relational query e

    Scalabig-datajavajdbc
    View on GitHub↗43,467
  • bytewax/bytewaxbytewax avatar

    bytewax/bytewax

    2,022View on GitHub↗

    Python Stream Processing

    Python
    View on GitHub↗2,022
  • apache/flinkapache avatar

    apache/flink

    26,086View on GitHub↗

    Apache Flink is a distributed processing engine designed for both high-throughput, low-latency data streams and finite batch workloads. It functions as a stateful stream processor and a SQL stream processing engine, providing a unified runtime to execute relational queries and event-based transformations. The system is distinguished by its ability to manage persistent operator state to ensure exactly-once processing guarantees and consistency during failures. It features specialized capabilities for complex event processing to detect temporal patterns and handles out-of-order events using eve

    Java
    View on GitHub↗26,086
  • cocoindex-io/cocoindexcocoindex-io avatar

    cocoindex-io/cocoindex

    6,117View on GitHub↗

    Cocoindex is an incremental data processing engine that builds and maintains live indexes for AI agents, with a core focus on codebase indexing and knowledge graph extraction. The engine uses a function-graph execution model where user-defined Python functions are composed into a directed acyclic graph, and it processes data incrementally so only changed source records or code paths are re-computed, avoiding full recomputation at any scale. It supports automatic schema inference from transformation pipeline type annotations and provides full data lineage tracing, tagging every output record wi

    Rustagentic-data-frameworkaiai-agents
    View on GitHub↗6,117
Compare all 30 related projects→

Frequently asked questions

What does apache/samza do?

Mirror of Apache Samza

What are the main features of apache/samza?

The main features of apache/samza are: Stream Processing, Streaming Engines.

Which projects share features with apache/samza?

Projects with overlapping indexed features include: cocoindex-io/cocoindex — Cocoindex is an incremental data processing engine that builds and maintains live indexes for AI agents, with a core… emqx/kuiper — Lightweight data stream processing engine for IoT edge. apache/spark — Apache Spark is a unified distributed data processing engine designed for large-scale data analysis and computation… bytewax/bytewax — Python Stream Processing. apache/flink — Apache Flink is a distributed processing engine designed for both high-throughput, low-latency data streams and finite… hstreamdb/hstream — HStreamDB is an open-source, cloud-native streaming database for IoT and beyond. Modernize your data stack for…