awesome-repositories.com
Blog
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectAboutHow we rankPressMCP server
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
apache avatar

apache/streampark

0
View on GitHub↗
4,312 stars·1,078 forks·Java·Apache-2.0·2 viewsstreampark.apache.org↗

Streampark

StreamPark is a centralized management platform designed to coordinate the deployment, monitoring, and operational lifecycle of distributed stream processing and batch applications. It functions as a control plane and orchestrator for data pipelines, specifically providing management capabilities for Apache Flink and Hadoop YARN environments.

The platform distinguishes itself through a low-code approach to task deployment and a multi-engine execution adapter that supports diverse processing runtimes. It facilitates real-time data pipeline management by combining streaming SQL analytics with a resource-based deployment pipeline that handles versioning, binary uploads, and savepoint-based state recovery.

The system covers a broad set of capabilities including distributed job orchestration, real-time data integration via pre-built connectors, and identity integration through LDAP or SSO. It also provides observability tools for second-level application monitoring and automated operational fault notifications.

Features

  • Distributed Job Orchestration - Provides a centralized framework for linking dependent tasks and managing the full lifecycle of distributed jobs.
  • Apache Flink Management Platforms - Provides a centralized control plane for deploying, monitoring, and managing Apache Flink applications.
  • Execution Engine Translation - Translates high-level task definitions into engine-specific configurations via an execution engine translation layer.
  • Unified Batch and Stream Processing Engines - Executes both real-time streaming and batch workloads across different versions of processing engines.
  • Low-Code Pipeline Managers - Offers a low-code platform for integrating data sources, executing streaming SQL, and managing warehouse connections.
  • Stream and Pipeline Orchestration - Coordinates the flow, transformation, and distributed processing of continuous data streams across multiple engine versions.
  • Distributed Task Schedulers - Orchestrates and distributes complex data processing workflows across computing clusters.
  • Real-Time Data Integration Platforms - Synchronizes live data across heterogeneous storage and analytical environments using pre-built connectors.
  • Streaming State Recovery - Recovers incremental operator states and output results from durable storage after failures.
  • Unified Web Interfaces - Provides a unified web interface for building, deploying, and operating streaming applications across various engines.
  • Centralized Application Maintenance Platforms - Coordinates the debugging, deployment, and maintenance of streaming applications through a centralized management platform.
  • Hub-and-Spoke Control Planes - Implements a centralized control plane to manage the deployment and monitoring of batch and stream tasks.
  • Stream Processing Coordinators - Coordinates the deployment and lifecycle of streaming and batch applications from a single centralized control point.
  • Centralized Control Planes - Provides a centralized control interface for deploying and operating real-time data pipelines across multiple engines.
  • Application Lifecycle Management - Manages the operational state, version updates, and health of running distributed stream processing tasks.
  • YARN and Cluster Management - Manages resources and executes streaming jobs within Hadoop YARN clusters.
  • Data Connectors - Utilizes pre-built connectors to efficiently acquire and move data between various sources and destinations.
  • Development Scaffolding - Provides standardized configurations and development scaffolds to implement streaming logic across multiple processing engines.
  • Data Processing Tasks - Handles compute-intensive operations including startup, savepoints, and performance analysis of stream jobs.
  • Data Warehouse Integrations - Links processing engines with open-source data lakes and real-time warehouses to streamline storage and retrieval.
  • Hadoop Workflow Orchestrators - Schedules and monitors big data workloads running on Hadoop clusters via HDFS.
  • Real-Time Data Streaming - Executes interactive streaming queries and manages streaming warehouses for real-time data analysis.
  • SQL Statement Executions - Provides capabilities to register tables and execute SQL statements for interfacing with external systems.
  • Streaming SQL Transformations - Executes SQL queries directly against live data streams for real-time filtering and aggregation.
  • Stream Processing Scaffolds - Provides development scaffolds and connectors to create stream processing applications compatible with multiple engine versions.
  • Job Resource Uploaders - Handles the uploading of binaries and configuration dependencies to distributed filesystems prior to job submission.
  • Low-Code Automation Platforms - Provides a low-code interface to coordinate the compilation, publishing, and deployment of processing tasks.
  • Task Release Tracking - Tracks custom files and task releases to facilitate coordinated updates and rollbacks across multiple versions.
  • Task Lifecycle Management - Governs the execution flow of tasks through explicit stages from initialization to destruction.
  • Metadata-Driven Orchestration - Uses an external database for metadata-driven orchestration of distributed processing job states and configurations.
  • Real-Time Application Performance Monitors - Provides live analysis of application performance metrics and health with second-level monitoring.
  • Stream Performance Monitoring - Tracks latency, throughput, and health of distributed data processing pipelines with automated alerts.

Star history

Star history chart for apache/streamparkStar history chart for apache/streampark

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to Streampark

Similar open-source projects, ranked by how many features they share with Streampark.
  • datalinkdc/dinkyDataLinkDC avatar

    DataLinkDC/dinky

    3,740View on GitHub↗

    Dinky is a real-time data platform for developing, deploying, and operating streaming applications based on Apache Flink. It functions as a SQL streaming IDE and a real-time data pipeline orchestrator, providing a web-based environment for writing and verifying queries with integrated logic plan visualization and lineage tracking. The platform acts as a distributed cluster manager, allowing the registration, monitoring, and administration of multiple processing clusters from a centralized interface. It also serves as a change data capture integration tool, synchronizing real-time database cha

    Javadatalakedatawarehouseflink
    View on GitHub↗3,740
  • hazelcast/hazelcasthazelcast avatar

    hazelcast/hazelcast

    6,570View on GitHub↗

    Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to support real-time analytics and event-driven applications. It functions as a partitioned, distributed key-value store that replicates data across cluster nodes to provide low-latency access and high availability. The platform also serves as a distributed SQL query engine, allowing users to execute standard SQL statements against both in-memory datasets and external data sources. What distinguishes Hazelcast is its use of a distributed consensus subsystem to maintain strongly consis

    Javabig-datacachingdata-in-motion
    View on GitHub↗6,570
  • zhisheng17/flink-learningzhisheng17 avatar

    zhisheng17/flink-learning

    15,071View on GitHub↗

    This project is a collection of educational resources and reference implementations for the Apache Flink stream processing framework. It provides a learning resource focused on mastering distributed stream processing through implementation guides, performance tuning tutorials, and practical examples. The repository features detailed walkthroughs for building real-time data pipelines using the DataStream and Table APIs. It includes specific integration examples for connecting Apache Flink with Kafka brokers and Elasticsearch indices, as well as reference implementations for real-time deduplica

    Javaclickhouseelasticsearchflink
    View on GitHub↗15,071
  • risingwavelabs/risingwaverisingwavelabs avatar

    risingwavelabs/risingwave

    9,093View on GitHub↗

    RisingWave is a cloud-native streaming database and real-time analytics engine that uses standard SQL to process continuous data streams. It functions as a streaming data lakehouse, combining the capabilities of a streaming SQL database with a platform that integrates streaming ingestion with open table formats. The system is distinguished by its use of the PostgreSQL wire protocol, allowing it to integrate with existing SQL tools and drivers. It employs a decoupled compute and storage architecture, persisting streaming state and materialized views in cloud object storage to enable independen

    Rustapache-icebergdata-engineeringdatabase
    View on GitHub↗9,093
See all 30 alternatives to Streampark→

Frequently asked questions

What does apache/streampark do?

StreamPark is a centralized management platform designed to coordinate the deployment, monitoring, and operational lifecycle of distributed stream processing and batch applications. It functions as a control plane and orchestrator for data pipelines, specifically providing management capabilities for Apache Flink and Hadoop YARN environments.

What are the main features of apache/streampark?

The main features of apache/streampark are: Distributed Job Orchestration, Apache Flink Management Platforms, Execution Engine Translation, Unified Batch and Stream Processing Engines, Low-Code Pipeline Managers, Stream and Pipeline Orchestration, Distributed Task Schedulers, Real-Time Data Integration Platforms.

What are some open-source alternatives to apache/streampark?

Open-source alternatives to apache/streampark include: datalinkdc/dinky — Dinky is a real-time data platform for developing, deploying, and operating streaming applications based on Apache… hazelcast/hazelcast — Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to… zhisheng17/flink-learning — This project is a collection of educational resources and reference implementations for the Apache Flink stream… risingwavelabs/risingwave — RisingWave is a cloud-native streaming database and real-time analytics engine that uses standard SQL to process… prefecthq/prefect — Prefect is a workflow orchestration platform designed to define, schedule, and monitor complex data pipelines as… apache/pinot — Pinot is a distributed, columnar analytical database designed for high-concurrency, low-latency query processing. It…