awesome-repositories.com
Blog
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoAcerca deCómo clasificamosPrensaServidor MCP
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
apache avatar

apache/streampark

0
View on GitHub↗
4,312 estrellas·1,078 forks·Java·Apache-2.0·2 vistasstreampark.apache.org↗

Streampark

StreamPark es una plataforma de gestión centralizada diseñada para coordinar el despliegue, monitoreo y ciclo de vida operativo de aplicaciones de procesamiento de flujos distribuidos y procesamiento por lotes (batch). Funciona como un plano de control y orquestador para pipelines de datos, proporcionando específicamente capacidades de gestión para entornos Apache Flink y Hadoop YARN.

La plataforma se distingue por un enfoque de bajo código para el despliegue de tareas y un adaptador de ejecución multi-motor que admite diversos runtimes de procesamiento. Facilita la gestión de pipelines de datos en tiempo real combinando análisis SQL de streaming con un pipeline de despliegue basado en recursos que maneja el versionado, subidas de binarios y recuperación de estado basada en savepoints.

El sistema cubre un amplio conjunto de capacidades, incluyendo orquestación de trabajos distribuidos, integración de datos en tiempo real a través de conectores preconstruidos e integración de identidad a través de LDAP o SSO. También proporciona herramientas de observabilidad para el monitoreo de aplicaciones de segundo nivel y notificaciones operativas automatizadas de fallos.

Features

  • Distributed Job Orchestration - Provides a centralized framework for linking dependent tasks and managing the full lifecycle of distributed jobs.
  • Apache Flink Management Platforms - Provides a centralized control plane for deploying, monitoring, and managing Apache Flink applications.
  • Execution Engine Translation - Translates high-level task definitions into engine-specific configurations via an execution engine translation layer.
  • Unified Batch and Stream Processing Engines - Executes both real-time streaming and batch workloads across different versions of processing engines.
  • Low-Code Pipeline Managers - Offers a low-code platform for integrating data sources, executing streaming SQL, and managing warehouse connections.
  • Stream and Pipeline Orchestration - Coordinates the flow, transformation, and distributed processing of continuous data streams across multiple engine versions.
  • Distributed Task Schedulers - Orchestrates and distributes complex data processing workflows across computing clusters.
  • Real-Time Data Integration Platforms - Synchronizes live data across heterogeneous storage and analytical environments using pre-built connectors.
  • Streaming State Recovery - Recovers incremental operator states and output results from durable storage after failures.
  • Unified Web Interfaces - Provides a unified web interface for building, deploying, and operating streaming applications across various engines.
  • Centralized Application Maintenance Platforms - Coordinates the debugging, deployment, and maintenance of streaming applications through a centralized management platform.
  • Hub-and-Spoke Control Planes - Implements a centralized control plane to manage the deployment and monitoring of batch and stream tasks.
  • Stream Processing Coordinators - Coordinates the deployment and lifecycle of streaming and batch applications from a single centralized control point.
  • Centralized Control Planes - Provides a centralized control interface for deploying and operating real-time data pipelines across multiple engines.
  • Application Lifecycle Management - Manages the operational state, version updates, and health of running distributed stream processing tasks.
  • YARN and Cluster Management - Manages resources and executes streaming jobs within Hadoop YARN clusters.
  • Data Connectors - Utilizes pre-built connectors to efficiently acquire and move data between various sources and destinations.
  • Development Scaffolding - Provides standardized configurations and development scaffolds to implement streaming logic across multiple processing engines.
  • Data Processing Tasks - Handles compute-intensive operations including startup, savepoints, and performance analysis of stream jobs.
  • Data Warehouse Integrations - Links processing engines with open-source data lakes and real-time warehouses to streamline storage and retrieval.
  • Hadoop Workflow Orchestrators - Schedules and monitors big data workloads running on Hadoop clusters via HDFS.
  • Real-Time Data Streaming - Executes interactive streaming queries and manages streaming warehouses for real-time data analysis.
  • SQL Statement Executions - Provides capabilities to register tables and execute SQL statements for interfacing with external systems.
  • Streaming SQL Transformations - Executes SQL queries directly against live data streams for real-time filtering and aggregation.
  • Stream Processing Scaffolds - Provides development scaffolds and connectors to create stream processing applications compatible with multiple engine versions.
  • Job Resource Uploaders - Handles the uploading of binaries and configuration dependencies to distributed filesystems prior to job submission.
  • Low-Code Automation Platforms - Provides a low-code interface to coordinate the compilation, publishing, and deployment of processing tasks.
  • Task Release Tracking - Tracks custom files and task releases to facilitate coordinated updates and rollbacks across multiple versions.
  • Task Lifecycle Management - Governs the execution flow of tasks through explicit stages from initialization to destruction.
  • Metadata-Driven Orchestration - Uses an external database for metadata-driven orchestration of distributed processing job states and configurations.
  • Real-Time Application Performance Monitors - Provides live analysis of application performance metrics and health with second-level monitoring.
  • Stream Performance Monitoring - Tracks latency, throughput, and health of distributed data processing pipelines with automated alerts.

Historial de estrellas

Gráfico del historial de estrellas de apache/streamparkGráfico del historial de estrellas de apache/streampark

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Preguntas frecuentes

¿Qué hace apache/streampark?

StreamPark es una plataforma de gestión centralizada diseñada para coordinar el despliegue, monitoreo y ciclo de vida operativo de aplicaciones de procesamiento de flujos distribuidos y procesamiento por lotes (batch). Funciona como un plano de control y orquestador para pipelines de datos, proporcionando específicamente capacidades de gestión para entornos Apache Flink y Hadoop YARN.

¿Cuáles son las características principales de apache/streampark?

Las características principales de apache/streampark son: Distributed Job Orchestration, Apache Flink Management Platforms, Execution Engine Translation, Unified Batch and Stream Processing Engines, Low-Code Pipeline Managers, Stream and Pipeline Orchestration, Distributed Task Schedulers, Real-Time Data Integration Platforms.

¿Qué alternativas de código abierto existen para apache/streampark?

Las alternativas de código abierto para apache/streampark incluyen: datalinkdc/dinky — Dinky is a real-time data platform for developing, deploying, and operating streaming applications based on Apache… hazelcast/hazelcast — Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to… zhisheng17/flink-learning — This project is a collection of educational resources and reference implementations for the Apache Flink stream… risingwavelabs/risingwave — RisingWave is a cloud-native streaming database and real-time analytics engine that uses standard SQL to process… prefecthq/prefect — Prefect is a workflow orchestration platform designed to define, schedule, and monitor complex data pipelines as… apache/pinot — Pinot is a distributed, columnar analytical database designed for high-concurrency, low-latency query processing. It…

Alternativas open-source a Streampark

Proyectos open-source similares, clasificados según cuántas características comparten con Streampark.
  • datalinkdc/dinkyAvatar de DataLinkDC

    DataLinkDC/dinky

    3,740Ver en GitHub↗

    Dinky is a real-time data platform for developing, deploying, and operating streaming applications based on Apache Flink. It functions as a SQL streaming IDE and a real-time data pipeline orchestrator, providing a web-based environment for writing and verifying queries with integrated logic plan visualization and lineage tracking. The platform acts as a distributed cluster manager, allowing the registration, monitoring, and administration of multiple processing clusters from a centralized interface. It also serves as a change data capture integration tool, synchronizing real-time database cha

    Javadatalakedatawarehouseflink
    Ver en GitHub↗3,740
  • hazelcast/hazelcastAvatar de hazelcast

    hazelcast/hazelcast

    6,570Ver en GitHub↗

    Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to support real-time analytics and event-driven applications. It functions as a partitioned, distributed key-value store that replicates data across cluster nodes to provide low-latency access and high availability. The platform also serves as a distributed SQL query engine, allowing users to execute standard SQL statements against both in-memory datasets and external data sources. What distinguishes Hazelcast is its use of a distributed consensus subsystem to maintain strongly consis

    Javabig-datacachingdata-in-motion
    Ver en GitHub↗6,570
  • zhisheng17/flink-learningAvatar de zhisheng17

    zhisheng17/flink-learning

    15,071Ver en GitHub↗

    This project is a collection of educational resources and reference implementations for the Apache Flink stream processing framework. It provides a learning resource focused on mastering distributed stream processing through implementation guides, performance tuning tutorials, and practical examples. The repository features detailed walkthroughs for building real-time data pipelines using the DataStream and Table APIs. It includes specific integration examples for connecting Apache Flink with Kafka brokers and Elasticsearch indices, as well as reference implementations for real-time deduplica

    Javaclickhouseelasticsearchflink
    Ver en GitHub↗15,071
  • risingwavelabs/risingwaveAvatar de risingwavelabs

    risingwavelabs/risingwave

    9,093Ver en GitHub↗

    RisingWave is a cloud-native streaming database and real-time analytics engine that uses standard SQL to process continuous data streams. It functions as a streaming data lakehouse, combining the capabilities of a streaming SQL database with a platform that integrates streaming ingestion with open table formats. The system is distinguished by its use of the PostgreSQL wire protocol, allowing it to integrate with existing SQL tools and drivers. It employs a decoupled compute and storage architecture, persisting streaming state and materialized views in cloud object storage to enable independen

    Rustapache-icebergdata-engineeringdatabase
    Ver en GitHub↗9,093
  • Ver las 30 alternativas a Streampark→