awesome-repositories.com
Blog
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
apache avatar

apache/streampark

0
View on GitHub↗
4,312 Stars·1,078 Forks·Java·Apache-2.0·2 Aufrufestreampark.apache.org↗

Streampark

StreamPark ist eine zentralisierte Managementplattform, die darauf ausgelegt ist, das Deployment, Monitoring und den operativen Lebenszyklus verteilter Stream-Processing- und Batch-Anwendungen zu koordinieren. Sie fungiert als Control-Plane und Orchestrator für Datenpipelines und bietet spezifisch Managementfunktionen für Apache Flink- und Hadoop YARN-Umgebungen.

Die Plattform zeichnet sich durch einen Low-Code-Ansatz für das Task-Deployment und einen Multi-Engine-Execution-Adapter aus, der diverse Verarbeitungs-Runtimes unterstützt. Sie erleichtert das Echtzeit-Datenpipeline-Management durch die Kombination von Streaming-SQL-Analytics mit einer ressourcenbasierten Deployment-Pipeline, die Versionierung, Binär-Uploads und Savepoint-basierte Zustands-Wiederherstellung handhabt.

Das System deckt ein breites Spektrum an Funktionen ab, einschließlich verteilter Job-Orchestrierung, Echtzeit-Datenintegration über vorgefertigte Connectors und Identitätsintegration via LDAP oder SSO. Es bietet zudem Observability-Tools für sekundengenaue Anwendungsüberwachung und automatisierte operative Fehlerbenachrichtigungen.

Features

  • Distributed Job Orchestration - Provides a centralized framework for linking dependent tasks and managing the full lifecycle of distributed jobs.
  • Apache Flink Management Platforms - Provides a centralized control plane for deploying, monitoring, and managing Apache Flink applications.
  • Execution Engine Translation - Translates high-level task definitions into engine-specific configurations via an execution engine translation layer.
  • Unified Batch and Stream Processing Engines - Executes both real-time streaming and batch workloads across different versions of processing engines.
  • Low-Code Pipeline Managers - Offers a low-code platform for integrating data sources, executing streaming SQL, and managing warehouse connections.
  • Stream and Pipeline Orchestration - Coordinates the flow, transformation, and distributed processing of continuous data streams across multiple engine versions.
  • Distributed Task Schedulers - Orchestrates and distributes complex data processing workflows across computing clusters.
  • Real-Time Data Integration Platforms - Synchronizes live data across heterogeneous storage and analytical environments using pre-built connectors.
  • Streaming State Recovery - Recovers incremental operator states and output results from durable storage after failures.
  • Unified Web Interfaces - Provides a unified web interface for building, deploying, and operating streaming applications across various engines.
  • Centralized Application Maintenance Platforms - Coordinates the debugging, deployment, and maintenance of streaming applications through a centralized management platform.
  • Hub-and-Spoke Control Planes - Implements a centralized control plane to manage the deployment and monitoring of batch and stream tasks.
  • Stream Processing Coordinators - Coordinates the deployment and lifecycle of streaming and batch applications from a single centralized control point.
  • Centralized Control Planes - Provides a centralized control interface for deploying and operating real-time data pipelines across multiple engines.
  • Application Lifecycle Management - Manages the operational state, version updates, and health of running distributed stream processing tasks.
  • YARN and Cluster Management - Manages resources and executes streaming jobs within Hadoop YARN clusters.
  • Data Connectors - Utilizes pre-built connectors to efficiently acquire and move data between various sources and destinations.
  • Development Scaffolding - Provides standardized configurations and development scaffolds to implement streaming logic across multiple processing engines.
  • Data Processing Tasks - Handles compute-intensive operations including startup, savepoints, and performance analysis of stream jobs.
  • Data Warehouse Integrations - Links processing engines with open-source data lakes and real-time warehouses to streamline storage and retrieval.
  • Hadoop Workflow Orchestrators - Schedules and monitors big data workloads running on Hadoop clusters via HDFS.
  • Real-Time Data Streaming - Executes interactive streaming queries and manages streaming warehouses for real-time data analysis.
  • SQL Statement Executions - Provides capabilities to register tables and execute SQL statements for interfacing with external systems.
  • Streaming SQL Transformations - Executes SQL queries directly against live data streams for real-time filtering and aggregation.
  • Stream Processing Scaffolds - Provides development scaffolds and connectors to create stream processing applications compatible with multiple engine versions.
  • Job Resource Uploaders - Handles the uploading of binaries and configuration dependencies to distributed filesystems prior to job submission.
  • Low-Code Automation Platforms - Provides a low-code interface to coordinate the compilation, publishing, and deployment of processing tasks.
  • Task Release Tracking - Tracks custom files and task releases to facilitate coordinated updates and rollbacks across multiple versions.
  • Task Lifecycle Management - Governs the execution flow of tasks through explicit stages from initialization to destruction.
  • Metadata-Driven Orchestration - Uses an external database for metadata-driven orchestration of distributed processing job states and configurations.
  • Real-Time Application Performance Monitors - Provides live analysis of application performance metrics and health with second-level monitoring.
  • Stream Performance Monitoring - Tracks latency, throughput, and health of distributed data processing pipelines with automated alerts.

Star-Verlauf

Star-Verlauf für apache/streamparkStar-Verlauf für apache/streampark

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Häufig gestellte Fragen

Was macht apache/streampark?

StreamPark ist eine zentralisierte Managementplattform, die darauf ausgelegt ist, das Deployment, Monitoring und den operativen Lebenszyklus verteilter Stream-Processing- und Batch-Anwendungen zu koordinieren. Sie fungiert als Control-Plane und Orchestrator für Datenpipelines und bietet spezifisch Managementfunktionen für Apache Flink- und Hadoop YARN-Umgebungen.

Was sind die Hauptfunktionen von apache/streampark?

Die Hauptfunktionen von apache/streampark sind: Distributed Job Orchestration, Apache Flink Management Platforms, Execution Engine Translation, Unified Batch and Stream Processing Engines, Low-Code Pipeline Managers, Stream and Pipeline Orchestration, Distributed Task Schedulers, Real-Time Data Integration Platforms.

Welche Open-Source-Alternativen gibt es zu apache/streampark?

Open-Source-Alternativen zu apache/streampark sind unter anderem: datalinkdc/dinky — Dinky is a real-time data platform for developing, deploying, and operating streaming applications based on Apache… hazelcast/hazelcast — Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to… zhisheng17/flink-learning — This project is a collection of educational resources and reference implementations for the Apache Flink stream… risingwavelabs/risingwave — RisingWave is a cloud-native streaming database and real-time analytics engine that uses standard SQL to process… prefecthq/prefect — Prefect is a workflow orchestration platform designed to define, schedule, and monitor complex data pipelines as… apache/pinot — Pinot is a distributed, columnar analytical database designed for high-concurrency, low-latency query processing. It…

Open-Source-Alternativen zu Streampark

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit Streampark.
  • datalinkdc/dinkyAvatar von DataLinkDC

    DataLinkDC/dinky

    3,740Auf GitHub ansehen↗

    Dinky is a real-time data platform for developing, deploying, and operating streaming applications based on Apache Flink. It functions as a SQL streaming IDE and a real-time data pipeline orchestrator, providing a web-based environment for writing and verifying queries with integrated logic plan visualization and lineage tracking. The platform acts as a distributed cluster manager, allowing the registration, monitoring, and administration of multiple processing clusters from a centralized interface. It also serves as a change data capture integration tool, synchronizing real-time database cha

    Javadatalakedatawarehouseflink
    Auf GitHub ansehen↗3,740
  • hazelcast/hazelcastAvatar von hazelcast

    hazelcast/hazelcast

    6,570Auf GitHub ansehen↗

    Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to support real-time analytics and event-driven applications. It functions as a partitioned, distributed key-value store that replicates data across cluster nodes to provide low-latency access and high availability. The platform also serves as a distributed SQL query engine, allowing users to execute standard SQL statements against both in-memory datasets and external data sources. What distinguishes Hazelcast is its use of a distributed consensus subsystem to maintain strongly consis

    Javabig-datacachingdata-in-motion
    Auf GitHub ansehen↗6,570
  • zhisheng17/flink-learningAvatar von zhisheng17

    zhisheng17/flink-learning

    15,071Auf GitHub ansehen↗

    This project is a collection of educational resources and reference implementations for the Apache Flink stream processing framework. It provides a learning resource focused on mastering distributed stream processing through implementation guides, performance tuning tutorials, and practical examples. The repository features detailed walkthroughs for building real-time data pipelines using the DataStream and Table APIs. It includes specific integration examples for connecting Apache Flink with Kafka brokers and Elasticsearch indices, as well as reference implementations for real-time deduplica

    Javaclickhouseelasticsearchflink
    Auf GitHub ansehen↗15,071
  • risingwavelabs/risingwaveAvatar von risingwavelabs

    risingwavelabs/risingwave

    9,093Auf GitHub ansehen↗

    RisingWave is a cloud-native streaming database and real-time analytics engine that uses standard SQL to process continuous data streams. It functions as a streaming data lakehouse, combining the capabilities of a streaming SQL database with a platform that integrates streaming ingestion with open table formats. The system is distinguished by its use of the PostgreSQL wire protocol, allowing it to integrate with existing SQL tools and drivers. It employs a decoupled compute and storage architecture, persisting streaming state and materialized views in cloud object storage to enable independen

    Rustapache-icebergdata-engineeringdatabase
    Auf GitHub ansehen↗9,093
  • Alle 30 Alternativen zu Streampark anzeigen→