awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

14 Repos

Awesome GitHub RepositoriesBig Data Frameworks

Infrastructure for distributed processing and large-scale data management.

Explore 14 awesome GitHub repositories matching part of an awesome list · Big Data Frameworks. Refine with filters or upvote what's useful.

Awesome Big Data Frameworks GitHub Repositories

Finde die besten Repos mit KI.Wir suchen mit KI nach den am besten passenden Repositories.
  • apache/flinkAvatar von apache

    apache/flink

    26,086Auf GitHub ansehen↗

    Apache Flink is a distributed processing engine designed for both high-throughput, low-latency data streams and finite batch workloads. It functions as a stateful stream processor and a SQL stream processing engine, providing a unified runtime to execute relational queries and event-based transformations. The system is distinguished by its ability to manage persistent operator state to ensure exactly-once processing guarantees and consistency during failures. It features specialized capabilities for complex event processing to detect temporal patterns and handles out-of-order events using eve

    Stream and batch processing framework for big data.

    Java
    Auf GitHub ansehen↗26,086
  • apache/hadoopAvatar von apache

    apache/hadoop

    15,567Auf GitHub ansehen↗

    Hadoop is a big data infrastructure suite and distributed data processing framework designed to store and process massive datasets across clusters of computers. It consists of a distributed storage system for managing large files across multiple nodes and a parallel computing engine for processing data across a distributed cluster. The framework implements a distributed file system to ensure fault tolerance and high throughput, paired with a programming model that processes large datasets in parallel. It manages the underlying hardware and software environment required for distributed big dat

    Framework for distributed processing of large datasets.

    Java
    Auf GitHub ansehen↗15,567
  • apache/stormAvatar von apache

    apache/storm

    6,683Auf GitHub ansehen↗

    Storm is a distributed stream processing framework designed to execute unbounded computations across a cluster to process real-time data streams. It functions as a data pipeline orchestrator that allows users to define and deploy declarative data flow graphs connecting streaming sources to processing components. The system operates as a multi-tenant distributed compute engine that isolates workloads and limits resource usage across shared clusters using dedicated pools and access control. It is also a secure distributed processing engine that employs encrypted node communication and SSL-secur

    Distributed real-time computation system.

    Java
    Auf GitHub ansehen↗6,683
  • alibaba/jstormAvatar von alibaba

    alibaba/jstorm

    3,877Auf GitHub ansehen↗

    jStorm is a distributed stream processing engine designed for executing low-latency computations on high-volume data streams using Apache Storm topologies. It functions as a real-time data analytics platform and distributed task orchestrator that manages complex data pipelines via directed acyclic graph execution. The system provides a scalable framework for data pipeline management, incorporating backpressure-aware flow control to regulate ingestion rates and dynamic resource allocation to adjust computing resources based on real-time demand. It maintains compatibility with Apache Storm conf

    Distributed and fault-tolerant real-time computation system.

    Java
    Auf GitHub ansehen↗3,877
  • twitter/heronAvatar von twitter

    twitter/heron

    3,632Auf GitHub ansehen↗

    Apache Heron (Incubating) is a realtime, distributed, fault-tolerant stream processing engine from Twitter

    Real-time analytics platform designed for high-scale processing.

    Java
    Auf GitHub ansehen↗3,632
  • linkedin/gobblinAvatar von linkedin

    linkedin/gobblin

    2,267Auf GitHub ansehen↗

    A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems.

    Universal data ingestion framework for Hadoop.

    Java
    Auf GitHub ansehen↗2,267
  • h2oai/h2o-2Avatar von h2oai

    h2oai/h2o-2

    2,253Auf GitHub ansehen↗

    Please visit https://github.com/h2oai/h2o-3 for latest H2O

    Statistical and machine learning runtime for big data.

    Java
    Auf GitHub ansehen↗2,253
  • oryxproject/oryxAvatar von oryxproject

    oryxproject/oryx

    1,783Auf GitHub ansehen↗

    Oryx 2: Lambda architecture on Apache Spark, Apache Kafka for real-time large scale machine learning

    Lambda architecture implementation for real-time machine learning.

    Java
    Auf GitHub ansehen↗1,783
  • twitter/elephant-birdAvatar von twitter

    twitter/elephant-bird

    1,133Auf GitHub ansehen↗

    Twitter's collection of LZO and Protocol Buffer-related Hadoop, Pig, Hive, and HBase code.

    Collection of Hadoop-related code for serialization and storage.

    Java
    Auf GitHub ansehen↗1,133
  • google/mr4cAvatar von google

    google/mr4c

    901Auf GitHub ansehen↗

    Introduction to the MR4C repo

    Framework for running native code within Hadoop execution.

    Java
    Auf GitHub ansehen↗901
  • etsy/oculusAvatar von etsy

    etsy/oculus

    707Auf GitHub ansehen↗

    The metric correlation component of Etsy's Kale system

    Component for anomaly correlation in large systems.

    Java
    Auf GitHub ansehen↗707
  • yahoo/samoaY

    yahoo/samoa

    0Auf GitHub ansehen↗

    Platform for mining big data streams.

    Auf GitHub ansehen↗0
  • linkedin/datafuL

    linkedin/datafu

    0Auf GitHub ansehen↗

    Collection of libraries for large-scale data processing.

    Auf GitHub ansehen↗0
  • cloudera/oryxC

    cloudera/oryx

    0Auf GitHub ansehen↗

    Infrastructure for large-scale machine learning and predictive analytics.

    Auf GitHub ansehen↗0
  1. Home
  2. Part of an Awesome List
  3. Databases & Data
  4. Big Data Frameworks