awesome-repositories.com
Blog
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectDespreCum realizăm clasamentulPresăServer MCP
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
apache avatar

apache/streampark

0
View on GitHub↗
4,312 stele·1,078 fork-uri·Java·Apache-2.0·2 vizualizăristreampark.apache.org↗

Streampark

StreamPark is a centralized management platform designed to coordinate the deployment, monitoring, and operational lifecycle of distributed stream processing and batch applications. It functions as a control plane and orchestrator for data pipelines, specifically providing management capabilities for Apache Flink and Hadoop YARN environments.

The platform distinguishes itself through a low-code approach to task deployment and a multi-engine execution adapter that supports diverse processing runtimes. It facilitates real-time data pipeline management by combining streaming SQL analytics with a resource-based deployment pipeline that handles versioning, binary uploads, and savepoint-based state recovery.

The system covers a broad set of capabilities including distributed job orchestration, real-time data integration via pre-built connectors, and identity integration through LDAP or SSO. It also provides observability tools for second-level application monitoring and automated operational fault notifications.

Features

  • Distributed Job Orchestration - Provides a centralized framework for linking dependent tasks and managing the full lifecycle of distributed jobs.
  • Apache Flink Management Platforms - Provides a centralized control plane for deploying, monitoring, and managing Apache Flink applications.
  • Execution Engine Translation - Translates high-level task definitions into engine-specific configurations via an execution engine translation layer.
  • Unified Batch and Stream Processing Engines - Executes both real-time streaming and batch workloads across different versions of processing engines.
  • Low-Code Pipeline Managers - Offers a low-code platform for integrating data sources, executing streaming SQL, and managing warehouse connections.
  • Stream and Pipeline Orchestration - Coordinates the flow, transformation, and distributed processing of continuous data streams across multiple engine versions.
  • Distributed Task Schedulers - Orchestrates and distributes complex data processing workflows across computing clusters.
  • Real-Time Data Integration Platforms - Synchronizes live data across heterogeneous storage and analytical environments using pre-built connectors.
  • Streaming State Recovery - Recovers incremental operator states and output results from durable storage after failures.
  • Unified Web Interfaces - Provides a unified web interface for building, deploying, and operating streaming applications across various engines.
  • Centralized Application Maintenance Platforms - Coordinates the debugging, deployment, and maintenance of streaming applications through a centralized management platform.
  • Hub-and-Spoke Control Planes - Implements a centralized control plane to manage the deployment and monitoring of batch and stream tasks.
  • Stream Processing Coordinators - Coordinates the deployment and lifecycle of streaming and batch applications from a single centralized control point.
  • Centralized Control Planes - Provides a centralized control interface for deploying and operating real-time data pipelines across multiple engines.
  • Application Lifecycle Management - Manages the operational state, version updates, and health of running distributed stream processing tasks.
  • YARN and Cluster Management - Manages resources and executes streaming jobs within Hadoop YARN clusters.
  • Data Connectors - Utilizes pre-built connectors to efficiently acquire and move data between various sources and destinations.
  • Development Scaffolding - Provides standardized configurations and development scaffolds to implement streaming logic across multiple processing engines.
  • Data Processing Tasks - Handles compute-intensive operations including startup, savepoints, and performance analysis of stream jobs.
  • Data Warehouse Integrations - Links processing engines with open-source data lakes and real-time warehouses to streamline storage and retrieval.
  • Hadoop Workflow Orchestrators - Schedules and monitors big data workloads running on Hadoop clusters via HDFS.
  • Real-Time Data Streaming - Executes interactive streaming queries and manages streaming warehouses for real-time data analysis.
  • SQL Statement Executions - Provides capabilities to register tables and execute SQL statements for interfacing with external systems.
  • Streaming SQL Transformations - Executes SQL queries directly against live data streams for real-time filtering and aggregation.
  • Stream Processing Scaffolds - Provides development scaffolds and connectors to create stream processing applications compatible with multiple engine versions.
  • Job Resource Uploaders - Handles the uploading of binaries and configuration dependencies to distributed filesystems prior to job submission.
  • Low-Code Automation Platforms - Provides a low-code interface to coordinate the compilation, publishing, and deployment of processing tasks.
  • Task Release Tracking - Tracks custom files and task releases to facilitate coordinated updates and rollbacks across multiple versions.
  • Task Lifecycle Management - Governs the execution flow of tasks through explicit stages from initialization to destruction.
  • Metadata-Driven Orchestration - Uses an external database for metadata-driven orchestration of distributed processing job states and configurations.
  • Real-Time Application Performance Monitors - Provides live analysis of application performance metrics and health with second-level monitoring.
  • Stream Performance Monitoring - Tracks latency, throughput, and health of distributed data processing pipelines with automated alerts.

Istoric stele

Graficul istoricului de stele pentru apache/streamparkGraficul istoricului de stele pentru apache/streampark

Căutare AI

Explorează mai multe repository-uri excelente

Descrie ce ai nevoie în limbaj simplu — AI-ul sortează mii de proiecte open source selectate în funcție de relevanță.

Start searching with AI

Alternative open-source pentru Streampark

Proiecte open-source similare, clasificate după numărul de funcționalități comune cu Streampark.
  • datalinkdc/dinkyAvatar DataLinkDC

    DataLinkDC/dinky

    3,740Vezi pe GitHub↗

    Dinky is a real-time data platform for developing, deploying, and operating streaming applications based on Apache Flink. It functions as a SQL streaming IDE and a real-time data pipeline orchestrator, providing a web-based environment for writing and verifying queries with integrated logic plan visualization and lineage tracking. The platform acts as a distributed cluster manager, allowing the registration, monitoring, and administration of multiple processing clusters from a centralized interface. It also serves as a change data capture integration tool, synchronizing real-time database cha

    Javadatalakedatawarehouseflink
    Vezi pe GitHub↗3,740
  • hazelcast/hazelcastAvatar hazelcast

    hazelcast/hazelcast

    6,570Vezi pe GitHub↗

    Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to support real-time analytics and event-driven applications. It functions as a partitioned, distributed key-value store that replicates data across cluster nodes to provide low-latency access and high availability. The platform also serves as a distributed SQL query engine, allowing users to execute standard SQL statements against both in-memory datasets and external data sources. What distinguishes Hazelcast is its use of a distributed consensus subsystem to maintain strongly consis

    Javabig-datacachingdata-in-motion
    Vezi pe GitHub↗6,570
  • zhisheng17/flink-learningAvatar zhisheng17

    zhisheng17/flink-learning

    15,071Vezi pe GitHub↗

    This project is a collection of educational resources and reference implementations for the Apache Flink stream processing framework. It provides a learning resource focused on mastering distributed stream processing through implementation guides, performance tuning tutorials, and practical examples. The repository features detailed walkthroughs for building real-time data pipelines using the DataStream and Table APIs. It includes specific integration examples for connecting Apache Flink with Kafka brokers and Elasticsearch indices, as well as reference implementations for real-time deduplica

    Javaclickhouseelasticsearchflink
    Vezi pe GitHub↗15,071
  • risingwavelabs/risingwaveAvatar risingwavelabs

    risingwavelabs/risingwave

    9,093Vezi pe GitHub↗

    RisingWave is a cloud-native streaming database and real-time analytics engine that uses standard SQL to process continuous data streams. It functions as a streaming data lakehouse, combining the capabilities of a streaming SQL database with a platform that integrates streaming ingestion with open table formats. The system is distinguished by its use of the PostgreSQL wire protocol, allowing it to integrate with existing SQL tools and drivers. It employs a decoupled compute and storage architecture, persisting streaming state and materialized views in cloud object storage to enable independen

    Rustapache-icebergdata-engineeringdatabase
    Vezi pe GitHub↗9,093
Vezi toate cele 30 alternative pentru Streampark→

Întrebări frecvente

Ce face apache/streampark?

StreamPark is a centralized management platform designed to coordinate the deployment, monitoring, and operational lifecycle of distributed stream processing and batch applications. It functions as a control plane and orchestrator for data pipelines, specifically providing management capabilities for Apache Flink and Hadoop YARN environments.

Care sunt principalele funcționalități ale apache/streampark?

Principalele funcționalități ale apache/streampark sunt: Distributed Job Orchestration, Apache Flink Management Platforms, Execution Engine Translation, Unified Batch and Stream Processing Engines, Low-Code Pipeline Managers, Stream and Pipeline Orchestration, Distributed Task Schedulers, Real-Time Data Integration Platforms.

Care sunt câteva alternative open-source pentru apache/streampark?

Alternativele open-source pentru apache/streampark includ: datalinkdc/dinky — Dinky is a real-time data platform for developing, deploying, and operating streaming applications based on Apache… hazelcast/hazelcast — Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to… zhisheng17/flink-learning — This project is a collection of educational resources and reference implementations for the Apache Flink stream… risingwavelabs/risingwave — RisingWave is a cloud-native streaming database and real-time analytics engine that uses standard SQL to process… prefecthq/prefect — Prefect is a workflow orchestration platform designed to define, schedule, and monitor complex data pipelines as… apache/pinot — Pinot is a distributed, columnar analytical database designed for high-concurrency, low-latency query processing. It…