awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

21 repositorios

Awesome GitHub RepositoriesBig Data

Frameworks for large-scale data processing and distributed systems.

Explore 21 awesome GitHub repositories matching part of an awesome list · Big Data. Refine with filters or upvote what's useful.

Awesome Big Data GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • awesomedata/awesome-public-datasetsAvatar de awesomedata

    awesomedata/awesome-public-datasets

    75,979Ver en GitHub↗

    This project is a community-maintained, open-access directory of high-quality public datasets. It serves as a centralized reference point for researchers, developers, and data scientists to locate reliable information sources across a wide spectrum of industries and scientific fields. By providing a structured index, the repository facilitates the discovery of data necessary for exploratory analysis, machine learning model training, and the development of data-intensive applications. The directory distinguishes itself through a lightweight, platform-agnostic approach to resource indexing that

    Listed in the “Big Data” section of the Awesome awesome list.

    aaron-swartzawesome-public-datasetsdatasets
    Ver en GitHub↗75,979
  • apache/sparkAvatar de apache

    apache/spark

    43,467Ver en GitHub↗

    Apache Spark is a unified distributed data processing engine designed for large-scale data analysis and computation graphs. It functions as a distributed machine learning framework, a graph processing system, a real-time stream processor, and a SQL analytics engine. The system enables the execution of distributed SQL querying, large-scale graph analysis, and real-time stream analytics across clusters of machines. It also provides a scalable environment for implementing machine learning algorithms and predictive model development on massive datasets. The engine incorporates relational query e

    Unified analytics engine for large-scale data.

    Scalabig-datajavajdbc
    Ver en GitHub↗43,467
  • apache/kafkaAvatar de apache

    apache/kafka

    32,846Ver en GitHub↗

    Kafka is a distributed event streaming platform designed for capturing, storing, and processing real-time data streams across interconnected nodes. It functions as a distributed commit log, providing a fault-tolerant storage mechanism that records state changes sequentially to ensure data consistency and durability across distributed environments. The platform distinguishes itself through a partitioned commit log architecture that enables horizontal scaling and parallel processing of data streams. It integrates a stream processing engine for continuous transformations and aggregations, while

    Distributed event streaming platform.

    Javakafkascala
    Ver en GitHub↗32,846
  • apache/flinkAvatar de apache

    apache/flink

    26,086Ver en GitHub↗

    Apache Flink is a distributed processing engine designed for both high-throughput, low-latency data streams and finite batch workloads. It functions as a stateful stream processor and a SQL stream processing engine, providing a unified runtime to execute relational queries and event-based transformations. The system is distinguished by its ability to manage persistent operator state to ensure exactly-once processing guarantees and consistency during failures. It features specialized capabilities for complex event processing to detect temporal patterns and handles out-of-order events using eve

    Distributed stream and batch processing engine.

    Java
    Ver en GitHub↗26,086
  • oxnr/awesome-bigdataAvatar de oxnr

    oxnr/awesome-bigdata

    14,454Ver en GitHub↗

    This project is a curated directory of software, frameworks, and educational resources designed for building, scaling, and maintaining distributed data processing and storage architectures. It serves as a comprehensive index for the distributed computing ecosystem, helping users identify the appropriate tools for managing large-scale information systems. The repository functions as a central hub for data engineering, offering categorized access to technologies that support batch and stream processing, machine learning, and interactive querying. By organizing these resources, it assists in the

    Listed in the “Big Data” section of the Awesome awesome list.

    awesomeawesome-listbigdata
    Ver en GitHub↗14,454
  • igorbarinov/awesome-data-engineeringAvatar de igorbarinov

    igorbarinov/awesome-data-engineering

    8,306Ver en GitHub↗

    Listed in the “Big Data” section of the Awesome awesome list.

    awesomeawesome-list
    Ver en GitHub↗8,306
  • briatte/awesome-network-analysisAvatar de briatte

    briatte/awesome-network-analysis

    4,067Ver en GitHub↗

    A curated list of awesome network analysis resources.

    Listed in the “Big Data” section of the Awesome awesome list.

    Rawesomeawesome-listcomplex-networks
    Ver en GitHub↗4,067
  • manuzhang/awesome-streamingAvatar de manuzhang

    manuzhang/awesome-streaming

    2,986Ver en GitHub↗

    a curated list of awesome streaming frameworks, applications, etc

    Listed in the “Big Data” section of the Awesome awesome list.

    awesomeawesome-listlist
    Ver en GitHub↗2,986
  • spotify/scioAvatar de spotify

    spotify/scio

    2,628Ver en GitHub↗

    A Scala API for Apache Beam and Google Cloud Dataflow.

    Scala API for Apache Beam and Dataflow.

    Scala
    Ver en GitHub↗2,628
  • looly/elasticsearch-definitive-guide-cnAvatar de looly

    looly/elasticsearch-definitive-guide-cn

    2,046Ver en GitHub↗

    Elasticsearch权威指南中文版

    Chinese translation of the Elasticsearch definitive guide.

    Shell
    Ver en GitHub↗2,046
  • awesome-spark/awesome-sparkAvatar de awesome-spark

    awesome-spark/awesome-spark

    1,882Ver en GitHub↗

    A curated list of awesome Apache Spark packages and resources.

    Listed in the “Big Data” section of the Awesome awesome list.

    Shellapache-sparkawesomepyspark
    Ver en GitHub↗1,882
  • akka/alpakka-kafkaAvatar de akka

    akka/alpakka-kafka

    1,421Ver en GitHub↗

    Alpakka Kafka connector - Alpakka is a Reactive Enterprise Integration library for Java and Scala, based on Reactive Streams and Akka.

    Reactive Kafka connector for Akka.

    Scala
    Ver en GitHub↗1,421
  • youngwookim/awesome-hadoopAvatar de youngwookim

    youngwookim/awesome-hadoop

    1,117Ver en GitHub↗

    A curated list of amazingly awesome Hadoop and Hadoop ecosystem resources

    Listed in the “Big Data” section of the Awesome awesome list.

    Ver en GitHub↗1,117
  • touk/nussknackerAvatar de TouK

    TouK/nussknacker

    731Ver en GitHub↗

    Low-code tool for automating actions on real time data | Stream processing for the users.

    Low-code tool for real-time data automation.

    Scala
    Ver en GitHub↗731
  • sduff/awesome-splunkAvatar de sduff

    sduff/awesome-splunk

    159Ver en GitHub↗

    A collection of awesome resources for Splunk

    Listed in the “Big Data” section of the Awesome awesome list.

    awesomeawesome-listsplunk
    Ver en GitHub↗159
  • galliaproject/gallia-coreAvatar de galliaproject

    galliaproject/gallia-core

    88Ver en GitHub↗

    A schema-aware Scala library for data transformation

    Schema-aware data transformation library.

    Scala
    Ver en GitHub↗88
  • ambster-public/awesome-qlikAvatar de ambster-public

    ambster-public/awesome-qlik

    78Ver en GitHub↗

    A curated list of awesome Qlik extensions and resources for Qlik Sense and QlikView

    Listed in the “Big Data” section of the Awesome awesome list.

    analyticsawesomeawesome-list
    Ver en GitHub↗78
  • openmole/gridscaleAvatar de openmole

    openmole/gridscale

    30Ver en GitHub↗

    Scala library for accessing various file, batch systems, job schedulers and grid middlewares.

    Access library for grid and batch systems.

    Scala
    Ver en GitHub↗30
  • elasticsearch/elasticsearch-definitive-guideE

    elasticsearch/elasticsearch-definitive-guide

    0Ver en GitHub↗

    The official guide for mastering the Elasticsearch search engine.

    Ver en GitHub↗0
  • waylau/apache-spark-tutorialW

    waylau/apache-spark-tutorial

    0Ver en GitHub↗

    Tutorials for processing big data with Apache Spark.

    Ver en GitHub↗0
Ant.12Siguiente
  1. Home
  2. Part of an Awesome List
  3. Databases & Data
  4. Big Data