awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

21 مستودعات

Awesome GitHub RepositoriesBig Data

Frameworks for large-scale data processing and distributed systems.

Explore 21 awesome GitHub repositories matching part of an awesome list · Big Data. Refine with filters or upvote what's useful.

Awesome Big Data GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • awesomedata/awesome-public-datasetsالصورة الرمزية لـ awesomedata

    awesomedata/awesome-public-datasets

    75,979عرض على GitHub↗

    This project is a community-maintained, open-access directory of high-quality public datasets. It serves as a centralized reference point for researchers, developers, and data scientists to locate reliable information sources across a wide spectrum of industries and scientific fields. By providing a structured index, the repository facilitates the discovery of data necessary for exploratory analysis, machine learning model training, and the development of data-intensive applications. The directory distinguishes itself through a lightweight, platform-agnostic approach to resource indexing that

    Listed in the “Big Data” section of the Awesome awesome list.

    aaron-swartzawesome-public-datasetsdatasets
    عرض على GitHub↗75,979
  • apache/sparkالصورة الرمزية لـ apache

    apache/spark

    43,467عرض على GitHub↗

    Apache Spark is a unified distributed data processing engine designed for large-scale data analysis and computation graphs. It functions as a distributed machine learning framework, a graph processing system, a real-time stream processor, and a SQL analytics engine. The system enables the execution of distributed SQL querying, large-scale graph analysis, and real-time stream analytics across clusters of machines. It also provides a scalable environment for implementing machine learning algorithms and predictive model development on massive datasets. The engine incorporates relational query e

    Unified analytics engine for large-scale data.

    Scalabig-datajavajdbc
    عرض على GitHub↗43,467
  • apache/kafkaالصورة الرمزية لـ apache

    apache/kafka

    32,846عرض على GitHub↗

    Kafka is a distributed event streaming platform designed for capturing, storing, and processing real-time data streams across interconnected nodes. It functions as a distributed commit log, providing a fault-tolerant storage mechanism that records state changes sequentially to ensure data consistency and durability across distributed environments. The platform distinguishes itself through a partitioned commit log architecture that enables horizontal scaling and parallel processing of data streams. It integrates a stream processing engine for continuous transformations and aggregations, while

    Distributed event streaming platform.

    Javakafkascala
    عرض على GitHub↗32,846
  • apache/flinkالصورة الرمزية لـ apache

    apache/flink

    26,086عرض على GitHub↗

    Apache Flink is a distributed processing engine designed for both high-throughput, low-latency data streams and finite batch workloads. It functions as a stateful stream processor and a SQL stream processing engine, providing a unified runtime to execute relational queries and event-based transformations. The system is distinguished by its ability to manage persistent operator state to ensure exactly-once processing guarantees and consistency during failures. It features specialized capabilities for complex event processing to detect temporal patterns and handles out-of-order events using eve

    Distributed stream and batch processing engine.

    Java
    عرض على GitHub↗26,086
  • oxnr/awesome-bigdataالصورة الرمزية لـ oxnr

    oxnr/awesome-bigdata

    14,454عرض على GitHub↗

    This project is a curated directory of software, frameworks, and educational resources designed for building, scaling, and maintaining distributed data processing and storage architectures. It serves as a comprehensive index for the distributed computing ecosystem, helping users identify the appropriate tools for managing large-scale information systems. The repository functions as a central hub for data engineering, offering categorized access to technologies that support batch and stream processing, machine learning, and interactive querying. By organizing these resources, it assists in the

    Listed in the “Big Data” section of the Awesome awesome list.

    awesomeawesome-listbigdata
    عرض على GitHub↗14,454
  • igorbarinov/awesome-data-engineeringالصورة الرمزية لـ igorbarinov

    igorbarinov/awesome-data-engineering

    8,306عرض على GitHub↗

    Listed in the “Big Data” section of the Awesome awesome list.

    awesomeawesome-list
    عرض على GitHub↗8,306
  • briatte/awesome-network-analysisالصورة الرمزية لـ briatte

    briatte/awesome-network-analysis

    4,067عرض على GitHub↗

    A curated list of awesome network analysis resources.

    Listed in the “Big Data” section of the Awesome awesome list.

    Rawesomeawesome-listcomplex-networks
    عرض على GitHub↗4,067
  • manuzhang/awesome-streamingالصورة الرمزية لـ manuzhang

    manuzhang/awesome-streaming

    2,986عرض على GitHub↗

    a curated list of awesome streaming frameworks, applications, etc

    Listed in the “Big Data” section of the Awesome awesome list.

    awesomeawesome-listlist
    عرض على GitHub↗2,986
  • spotify/scioالصورة الرمزية لـ spotify

    spotify/scio

    2,628عرض على GitHub↗

    A Scala API for Apache Beam and Google Cloud Dataflow.

    Scala API for Apache Beam and Dataflow.

    Scala
    عرض على GitHub↗2,628
  • looly/elasticsearch-definitive-guide-cnالصورة الرمزية لـ looly

    looly/elasticsearch-definitive-guide-cn

    2,046عرض على GitHub↗

    Elasticsearch权威指南中文版

    Chinese translation of the Elasticsearch definitive guide.

    Shell
    عرض على GitHub↗2,046
  • awesome-spark/awesome-sparkالصورة الرمزية لـ awesome-spark

    awesome-spark/awesome-spark

    1,882عرض على GitHub↗

    A curated list of awesome Apache Spark packages and resources.

    Listed in the “Big Data” section of the Awesome awesome list.

    Shellapache-sparkawesomepyspark
    عرض على GitHub↗1,882
  • akka/alpakka-kafkaالصورة الرمزية لـ akka

    akka/alpakka-kafka

    1,421عرض على GitHub↗

    Alpakka Kafka connector - Alpakka is a Reactive Enterprise Integration library for Java and Scala, based on Reactive Streams and Akka.

    Reactive Kafka connector for Akka.

    Scala
    عرض على GitHub↗1,421
  • youngwookim/awesome-hadoopالصورة الرمزية لـ youngwookim

    youngwookim/awesome-hadoop

    1,117عرض على GitHub↗

    A curated list of amazingly awesome Hadoop and Hadoop ecosystem resources

    Listed in the “Big Data” section of the Awesome awesome list.

    عرض على GitHub↗1,117
  • touk/nussknackerالصورة الرمزية لـ TouK

    TouK/nussknacker

    731عرض على GitHub↗

    Low-code tool for automating actions on real time data | Stream processing for the users.

    Low-code tool for real-time data automation.

    Scala
    عرض على GitHub↗731
  • sduff/awesome-splunkالصورة الرمزية لـ sduff

    sduff/awesome-splunk

    159عرض على GitHub↗

    A collection of awesome resources for Splunk

    Listed in the “Big Data” section of the Awesome awesome list.

    awesomeawesome-listsplunk
    عرض على GitHub↗159
  • galliaproject/gallia-coreالصورة الرمزية لـ galliaproject

    galliaproject/gallia-core

    88عرض على GitHub↗

    A schema-aware Scala library for data transformation

    Schema-aware data transformation library.

    Scala
    عرض على GitHub↗88
  • ambster-public/awesome-qlikالصورة الرمزية لـ ambster-public

    ambster-public/awesome-qlik

    78عرض على GitHub↗

    A curated list of awesome Qlik extensions and resources for Qlik Sense and QlikView

    Listed in the “Big Data” section of the Awesome awesome list.

    analyticsawesomeawesome-list
    عرض على GitHub↗78
  • openmole/gridscaleالصورة الرمزية لـ openmole

    openmole/gridscale

    30عرض على GitHub↗

    Scala library for accessing various file, batch systems, job schedulers and grid middlewares.

    Access library for grid and batch systems.

    Scala
    عرض على GitHub↗30
  • elasticsearch/elasticsearch-definitive-guideE

    elasticsearch/elasticsearch-definitive-guide

    0عرض على GitHub↗

    The official guide for mastering the Elasticsearch search engine.

    عرض على GitHub↗0
  • waylau/apache-spark-tutorialW

    waylau/apache-spark-tutorial

    0عرض على GitHub↗

    Tutorials for processing big data with Apache Spark.

    عرض على GitHub↗0
السابق12التالي
  1. Home
  2. Part of an Awesome List
  3. Databases & Data
  4. Big Data