awesome-repositories.com
Blog
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
DataLinkDC avatar

DataLinkDC/dinky

0
View on GitHub↗
3,740 Stars·1,345 Forks·Java·Apache-2.0·4 Aufrufewww.dinky.org.cn↗

Dinky

Dinky is a real-time data platform for developing, deploying, and operating streaming applications based on Apache Flink. It functions as a SQL streaming IDE and a real-time data pipeline orchestrator, providing a web-based environment for writing and verifying queries with integrated logic plan visualization and lineage tracking.

The platform acts as a distributed cluster manager, allowing the registration, monitoring, and administration of multiple processing clusters from a centralized interface. It also serves as a change data capture integration tool, synchronizing real-time database changes into warehouses or lakes with automatic schema evolution.

The system covers a broad range of operational capabilities, including data pipeline observability via lineage analysis, enterprise data governance through multi-tenant resource isolation and role-based access control, and fault tolerance using snapshot-based state recovery. Developer experience is supported through custom function management, query result previewing, and a managed editor for SQL application authoring.

Operational oversight is provided through job health monitoring, real-time alert configurations, and execution plan inspection.

Features

  • Apache Flink Management Platforms - Acts as a centralized control plane for managing the full lifecycle of Apache Flink streaming and batch applications.
  • SQL Query Execution - Executes SQL queries across different deployment environments using a standardized interface to process data at scale.
  • SQL Development Environments - Provides a web-based integrated environment for writing, verifying, and optimizing SQL streaming applications.
  • Unified Metadata Catalogs - Provides a centralized interface to manage schemas and connection details across diverse external database engines.
  • Streaming SQL IDEs - Provides a web-based IDE for authoring streaming SQL with integrated logic plan visualization and lineage tracking.
  • Change Data Capture - Synchronizes databases in real-time to warehouses or lakes with automatic table creation and schema evolution.
  • Flink Execution Engines - Leverages Apache Flink as a distributed processing engine to execute real-time data pipelines and SQL jobs.
  • Data Lake Orchestrators - Orchestrates the flow of real-time data from various sources into cloud-native data lakes and warehouses.
  • Logic Verification - Provides real-time previewing of table contents and custom functions to verify streaming processing logic.
  • CDC Synchronization - Uses change data capture to synchronize real-time database changes into warehouses and lakes.
  • External Storage Integrations - Enables linking and importing diverse external database engines and storage systems for query development.
  • Multi-Tenant Resource Isolation - Organizes users and workloads into namespaces with role-based access control to secure shared cluster environments.
  • Real-Time Data Integration Platforms - Provides a platform for synchronizing live data between heterogeneous databases and analytical warehouses with automatic schema evolution.
  • Real-time Data Synchronization - Captures database changes in real-time to automate table creation and schema evolution in target warehouses.
  • Streaming Pipeline Orchestrators - Orchestrates the complete lifecycle of streaming jobs, covering deployment, external synchronization, and state recovery.
  • Streaming State Recovery - Restores streaming applications to specific historical points by managing checkpoints and savepoints within the cluster.
  • Streaming Application Authoring - Ships a lightweight code editor with completion and logic checking specifically for authoring data streaming jobs.
  • Centralized Environment Management - Registers and controls multiple remote processing clusters from a single centralized interface.
  • Distributed Cluster Management - Provides a centralized interface for the administration and orchestration of shared computing resources across multiple processing clusters.
  • Cluster Administration - Organizes configurations, data sources, and global variables within a centralized cluster management interface.
  • Streaming Cluster Orchestration - Orchestrates the full lifecycle of distributed streaming clusters, including deployment and resource administration.
  • Job Lifecycle Management - Provides full lifecycle controls for streaming tasks, including triggering starts, canceling processes, and managing recovery points.
  • Data Pipeline Lineage Inspectors - Tracks data lineage and visualizes transformation logic to provide deep observability into data pipelines.
  • SQL Lineage Analyzers - Analyzes SQL statements to track column-level data lineage and visualize transformation logic.
  • Lineage Mapping - Tracks the flow of data through applications to visualize transformations and analyze the impact of changes.
  • Change Filters - Excludes specific database operations from the synchronization stream using integrated configuration parameters.
  • Full Instance Synchronization - Builds real-time data pipelines to synchronize all tables from a complete database instance to downstream systems.
  • Data Lineage Visualizations - Maps dependencies between data elements to analyze the global impact of changes across the pipeline.
  • Stream Snapshot Restorers - Includes utilities to restore streaming application configurations and message history from stream-specific snapshots.
  • Multi-Destination Data Routing - Provides mechanisms for directing processed data streams to multiple external output targets and clusters.
  • Pluggable Connector Frameworks - Integrates diverse data sources and destinations through a standardized, pluggable connector architecture.
  • Source-to-Sink Table Mappings - Implements rules for matching and renaming source tables to destination tables using prefixes and suffixes.
  • Automated Sink Schema Generation - Automatically constructs destination table definitions by extracting metadata from source stages during synchronization.
  • Custom SQL Functions - Allows development and submission of custom SQL functions to handle specialized data processing tasks.
  • User-Defined Function Development - Enables writing and submitting custom processing logic using common programming languages to extend data processing.
  • SQL Schema Generators - Automatically derives destination SQL table and type definitions by capturing metadata from the source stage.
  • Streaming State Recovery - Recovers incremental operator states and results from durable storage to ensure streaming stability after failures.
  • Stream Processing Job Submissions - Submits stream processing pipelines to clusters using SQL statements and a managed interface.
  • Cluster Job Launchers - Dispatches developed streaming applications to remote compute resources on production clusters for persistent execution.
  • External Cluster Imports - Connects external processing clusters to a management platform by specifying names and addresses.
  • SQL Project Deployments - Deploys declarative SQL projects and tasks across diverse environments and session modes.
  • Streaming Job Execution - Deploys streaming tasks across multiple session or application modes in diverse environments.
  • Stream Processing Pipeline Deployments - Deploys and monitors fault-tolerant stream processing applications across local and distributed environments.
  • Enterprise Data Governance - Controls system access through multi-tenancy, role-based permissions, and centralized identity authentication.
  • Enterprise Security Integrations - Integrates identity verification protocols to secure access to resource centers for enterprise-grade security.
  • Multi-Tenant Administration - Provides administrative management of isolated namespaces, roles, and multi-tenancy configurations.
  • Role-Based Access Controls - Implements role-based access controls and multi-tenant isolation to manage system usage and permissions.
  • Single Sign-On Integrations - Centralizes user authentication through single sign-on integrations for access across multiple services.
  • Stateful Recovery Mechanisms - Implements stateful recovery mechanisms using snapshots and checkpoints to restore streaming jobs without data loss.
  • Execution Configurations - Defines execution modes and parallelism within a managed editor to configure streaming applications.
  • Real-Time Application Performance Monitors - Provides live analysis of streaming application performance and health via real-time monitoring views.
  • Cluster Health Monitoring - Provides interfaces for querying the status and health metrics of distributed processing clusters.
  • Catalog Management - Provides a unified interface to configure and manage multiple independent database catalogs for simultaneous querying.
  • Cluster Stability Monitoring - Monitors runtime information and cluster logs while managing snapshot recovery to maintain system stability.
  • Job Execution Tracking - Retrieves operational data and execution plans to track the lifecycle and processing logic of computational tasks.
  • Real-Time Application Log Monitoring - Implements real-time log monitoring and online controls to manage the operational lifecycle of streaming jobs.
  • Real-Time Monitoring - Observes live cluster health and exception logs in real-time to maintain streaming stability.
  • Query Result Previews - Runs select statements in a local environment to preview data output and verify logic before deployment.

Star-Verlauf

Star-Verlauf für datalinkdc/dinkyStar-Verlauf für datalinkdc/dinky

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Open-Source-Alternativen zu Dinky

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit Dinky.
  • hazelcast/hazelcastAvatar von hazelcast

    hazelcast/hazelcast

    6,570Auf GitHub ansehen↗

    Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to support real-time analytics and event-driven applications. It functions as a partitioned, distributed key-value store that replicates data across cluster nodes to provide low-latency access and high availability. The platform also serves as a distributed SQL query engine, allowing users to execute standard SQL statements against both in-memory datasets and external data sources. What distinguishes Hazelcast is its use of a distributed consensus subsystem to maintain strongly consis

    Javabig-datacachingdata-in-motion
    Auf GitHub ansehen↗6,570
  • zeebe-io/zeebeAvatar von zeebe-io

    zeebe-io/zeebe

    4,171Auf GitHub ansehen↗

    Zeebe is a cloud-native workflow engine and distributed state machine designed for business process orchestration using BPMN and DMN standards. It operates as a high-performance gRPC workflow runtime that executes complex business processes through a partitioned event-streaming architecture. The system also functions as an orchestrator for large language model agents, coordinating AI reasoning and tool use within deterministic business processes. The engine is distinguished by its peer-to-peer broker networking and a consensus-based data replication model that ensures high availability and fa

    Java
    Auf GitHub ansehen↗4,171
  • apache/pinotAvatar von apache

    apache/pinot

    6,098Auf GitHub ansehen↗

    Pinot is a distributed, columnar analytical database designed for high-concurrency, low-latency query processing. It functions as a real-time OLAP datastore, enabling interactive, user-facing analytics by ingesting and querying massive datasets from both streaming and batch sources. The system architecture relies on a centralized controller for cluster coordination and a distributed segment-based storage model to ensure horizontal scalability. The platform distinguishes itself through a hybrid ingestion pipeline that unifies real-time event streams and historical batch data into a single quer

    Java
    Auf GitHub ansehen↗6,098
  • apache/streamparkAvatar von apache

    apache/streampark

    4,312Auf GitHub ansehen↗

    StreamPark is a centralized management platform designed to coordinate the deployment, monitoring, and operational lifecycle of distributed stream processing and batch applications. It functions as a control plane and orchestrator for data pipelines, specifically providing management capabilities for Apache Flink and Hadoop YARN environments. The platform distinguishes itself through a low-code approach to task deployment and a multi-engine execution adapter that supports diverse processing runtimes. It facilitates real-time data pipeline management by combining streaming SQL analytics with a

    Javaapachedevelopment-frameworkeasy-to-use
    Auf GitHub ansehen↗4,312
Alle 30 Alternativen zu Dinky anzeigen→

Häufig gestellte Fragen

Was macht datalinkdc/dinky?

Dinky is a real-time data platform for developing, deploying, and operating streaming applications based on Apache Flink. It functions as a SQL streaming IDE and a real-time data pipeline orchestrator, providing a web-based environment for writing and verifying queries with integrated logic plan visualization and lineage tracking.

Was sind die Hauptfunktionen von datalinkdc/dinky?

Die Hauptfunktionen von datalinkdc/dinky sind: Apache Flink Management Platforms, SQL Query Execution, SQL Development Environments, Unified Metadata Catalogs, Streaming SQL IDEs, Change Data Capture, Flink Execution Engines, Data Lake Orchestrators.

Welche Open-Source-Alternativen gibt es zu datalinkdc/dinky?

Open-Source-Alternativen zu datalinkdc/dinky sind unter anderem: hazelcast/hazelcast — Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to… zeebe-io/zeebe — Zeebe is a cloud-native workflow engine and distributed state machine designed for business process orchestration… apache/pinot — Pinot is a distributed, columnar analytical database designed for high-concurrency, low-latency query processing. It… apache/streampark — StreamPark is a centralized management platform designed to coordinate the deployment, monitoring, and operational… apache/flink-cdc — This project is a streaming data integration framework that captures real-time database changes and synchronizes them… redis/redisinsight — RedisInsight is a graphical user interface and management tool for browsing, analyzing, and administering Redis…