6 مستودعات
Management of internal state for complex streaming operations like anti-joins and dynamic filtering.
Distinct from Query State Management: Specific to the low-latency state maintenance required for continuous streaming queries, unlike general query result state.
Explore 6 awesome GitHub repositories matching data & databases · Streaming State Management. Refine with filters or upvote what's useful.
RisingWave is a cloud-native streaming database and real-time analytics engine that uses standard SQL to process continuous data streams. It functions as a streaming data lakehouse, combining the capabilities of a streaming SQL database with a platform that integrates streaming ingestion with open table formats. The system is distinguished by its use of the PostgreSQL wire protocol, allowing it to integrate with existing SQL tools and drivers. It employs a decoupled compute and storage architecture, persisting streaming state and materialized views in cloud object storage to enable independen
Maintains low-latency state for complex streaming operations including anti-joins and dynamic filtering.
Storm is a distributed stream processing framework and fault-tolerant compute engine designed for executing real-time continuous computations across a cluster of machines. It functions as a stateful stream processor and cluster topology manager, enabling the deployment and monitoring of distributed data flow configurations. The system ensures exactly-once semantics by utilizing transactional state management to guarantee that every message in a data stream is processed exactly one time. It further operates as a distributed RPC system, allowing for the integration of non-native languages throu
Combines high-volume stream processing with distributed queries to maintain and retrieve real-time state.
Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to support real-time analytics and event-driven applications. It functions as a partitioned, distributed key-value store that replicates data across cluster nodes to provide low-latency access and high availability. The platform also serves as a distributed SQL query engine, allowing users to execute standard SQL statements against both in-memory datasets and external data sources. What distinguishes Hazelcast is its use of a distributed consensus subsystem to maintain strongly consis
Manages internal state for complex streaming operations including TTL-based eviction and explicit deletion.
Arroyo is a high-performance stream processing platform built in Rust. It executes continuous SQL queries on streaming data with event-time semantics, enabling accurate windowed aggregations, joins, and stateful computations on unbounded event streams. The platform uses native Rust execution for high throughput and low latency, with periodic checkpointing for exactly-once fault tolerance and horizontal scaling across distributed workers. The system integrates deeply with Kafka for reading and writing topics with exactly-once delivery and supports change data capture (CDC) from MySQL and Postg
Maintains state across streaming events to enable windowed aggregations, joins, and other stateful computations.
more-itertools is a Python iterable utility library providing advanced functions for manipulating, filtering, and transforming data sequences. It serves as a data stream processing toolkit and a set of utilities for iterator state management, extending the capabilities of the standard Python itertools module. The library includes a combinatorial math toolkit for generating permutations, combinations, and powersets, alongside routines for number theory calculations and matrix operations. It also provides tools for stream state management, allowing users to peek at upcoming elements or seek wit
Implements stateful wrappers that allow users to peek at future elements or seek within a sequence.
more-itertools هي مكتبة إضافية لوحدة itertools في Python. تعمل كمجموعة أدوات لمعالجة التكرارات، وتوفر مجموعة واسعة من الروتينات لتحويل البيانات، والتوليد التوافقي، وإدارة حالة المكرر. تتميز المكتبة بإدارة الحالة المتقدمة وتوليد التسلسل المعقد. توفر إمكانيات لإلقاء نظرة خاطفة على العناصر المستقبلية، والبحث داخل التسلسلات، وإنتاج التباديل الفريدة، والتوليفات، وتقسيمات المجموعات من المجموعات التي قد تحتوي على عناصر مكررة. يغطي سطح قدرتها الأوسع مهام معالجة البيانات مثل التسطيح العودي، والتجميع، والحشو، وإعادة تشكيل تدفقات البيانات. كما تتضمن أدوات لدمج التدفق، والنافذة لتحليل الحي المحلي، ومزامنة التكرار الآمن للخيوط. يوفر المشروع أيضاً روتينات متخصصة لمعالجة التسلسل الرقمي، بما في ذلك ضرب المصفوفات، والالتفاف الخطي المنفصل، وتحويلات Fourier.
Provides mechanisms for looking ahead at upcoming elements without consuming the iterator.