awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
oxnr avatar

oxnr/awesome-bigdata

0
View on GitHub↗
14,454 نجوم·2,582 تفرعات·MIT·7 مشاهداتgithub.com/onurakpolat/awesome-bigdata↗

Awesome Bigdata

This project is a curated directory of software, frameworks, and educational resources designed for building, scaling, and maintaining distributed data processing and storage architectures. It serves as a comprehensive index for the distributed computing ecosystem, helping users identify the appropriate tools for managing large-scale information systems.

The repository functions as a central hub for data engineering, offering categorized access to technologies that support batch and stream processing, machine learning, and interactive querying. By organizing these resources, it assists in the design and development of complex data pipelines and the selection of infrastructure components for massive datasets.

Features

  • Curated Software Directories - Acts as a structured directory of high-quality software projects and frameworks organized by technical domain for distributed data architectures.
  • Data Analytics Engines - Provides a comprehensive index of high-performance computational engines for executing complex analytical queries on massive datasets.
  • Awesome List - A community-curated directory that catalogs and links out to other open-source projects, rather than a standalone tool you run yourself.
  • Large-Scale Training Frameworks - Curates infrastructure and orchestration frameworks for scaling machine learning model training across massive compute clusters.
  • Data Discovery Tools - Helps identify and evaluate databases and processing frameworks for large-scale data infrastructure.
  • Data Pipeline Orchestration - Schedules recurring batch jobs and resolves task dependencies for reliable data pipeline execution.
  • Distributed Computing - Executes data processing tasks across interconnected nodes to handle massive datasets through parallel computation.
  • Developer Resource Hubs - Serves as a centralized hub for accessing educational materials, libraries, and tools for designing large-scale data pipelines.
  • Columnar Storage Engines - Organizes data into columns to optimize analytical query performance and compression on massive datasets.
  • Distributed Computing Engines - Indexes a wide range of distributed computing engines and frameworks for batch, stream, and interactive data processing.
  • Stream Processing Systems - Processes continuous data flows through distributed buffers for real-time analysis and state updates.
  • Distributed Data Processing - Executes batch and real-time data workflows across computing clusters using parallel programming models.
  • High-Performance Data Infrastructures - Maintains scalable, high-performance storage systems for structured and unstructured data across cloud environments.
  • Log-Structured Storage - Uses append-only file structures to optimize write throughput and retrieval in distributed storage systems.
  • Query Languages - Facilitates interactive analysis of large datasets using standard query languages.
  • Engineering Guides - Offers educational resources and libraries for designing and maintaining distributed data processing architectures.
  • Directed Acyclic Graph Engines - Manages complex data pipeline dependencies and execution order across distributed environments.
  • Machine Learning Training - Supports training predictive models on large distributed datasets to improve analytical accuracy.
  • Big Data - Listed in the “Big Data” section of the Awesome awesome list.
  • Awesome Lists - Resources for big data frameworks and tools.
  • Data Ingestion - Collects and synchronizes streaming data from external sources into storage or processing systems.
  • Data Visualization Dashboards - Enables the creation of interactive dashboards and charts to visualize complex data insights.
  • Access Control - Implements authentication and authorization controls to protect sensitive data within distributed environments.
  • Role-Based Access Control - Enforces security by mapping user identities to specific permissions for data and infrastructure access.
  • Query Optimization Engines - Translates high-level analytical requests into optimized execution plans for distributed processing engines.

سجل النجوم

مخطط تاريخ النجوم لـ oxnr/awesome-bigdataمخطط تاريخ النجوم لـ oxnr/awesome-bigdata

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

بدائل مفتوحة المصدر لـ Awesome Bigdata

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع Awesome Bigdata.
  • awesome-selfhosted/awesome-selfhostedالصورة الرمزية لـ awesome-selfhosted

    awesome-selfhosted/awesome-selfhosted

    299,516عرض على GitHub↗

    This project is a community-curated directory of open-source software designed for deployment in private server environments and home labs. It serves as a comprehensive resource for discovering independent, self-hosted alternatives to mainstream cloud services, enabling users to maintain full data ownership and control over their digital infrastructure. The directory is structured through a hierarchical taxonomy that organizes a vast collection of applications into logical categories, ranging from media management and data analytics to private communication and team productivity tools. It dis

    awesomeawesome-listcloud
    عرض على GitHub↗299,516
  • cjbarber/toolsofthetradeالصورة الرمزية لـ cjbarber

    cjbarber/ToolsOfTheTrade

    17,075عرض على GitHub↗

    ToolsOfTheTrade is a comprehensive, curated directory designed to help developers and operational teams discover software services for a wide range of technical and business requirements. It functions as a centralized resource catalog, indexing tools that support the entire software development lifecycle, infrastructure management, and organizational productivity. The project distinguishes itself by providing a structured index of both hosted and self-hosted solutions, enabling users to identify specific utilities for tasks ranging from continuous integration and project planning to applicati

    عرض على GitHub↗17,075
  • vonng/ddiaالصورة الرمزية لـ Vonng

    Vonng/ddia

    22,648عرض على GitHub↗

    This project serves as a comprehensive technical reference for the architecture and design of data-intensive applications. It provides a structured analysis of the fundamental principles required to build reliable, scalable, and maintainable software systems, covering the core trade-offs inherent in modern data infrastructure. The repository explores the mechanics of distributed data management, including strategies for replication, partitioning, and achieving consensus across multiple nodes. It details the design of storage engines, indexing techniques, and transaction management models, whi

    Pythonbookdatabaseddia
    عرض على GitHub↗22,648
  • hazelcast/hazelcastالصورة الرمزية لـ hazelcast

    hazelcast/hazelcast

    6,570عرض على GitHub↗

    Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to support real-time analytics and event-driven applications. It functions as a partitioned, distributed key-value store that replicates data across cluster nodes to provide low-latency access and high availability. The platform also serves as a distributed SQL query engine, allowing users to execute standard SQL statements against both in-memory datasets and external data sources. What distinguishes Hazelcast is its use of a distributed consensus subsystem to maintain strongly consis

    Javabig-datacachingdata-in-motion
    عرض على GitHub↗6,570
عرض جميع البدائل الـ 30 لـ Awesome Bigdata→

الأسئلة الشائعة

ما هي وظيفة oxnr/awesome-bigdata؟

This project is a curated directory of software, frameworks, and educational resources designed for building, scaling, and maintaining distributed data processing and storage architectures. It serves as a comprehensive index for the distributed computing ecosystem, helping users identify the appropriate tools for managing large-scale information systems.

ما هي الميزات الرئيسية لـ oxnr/awesome-bigdata؟

الميزات الرئيسية لـ oxnr/awesome-bigdata هي: Curated Software Directories, Data Analytics Engines, Awesome List, Large-Scale Training Frameworks, Data Discovery Tools, Data Pipeline Orchestration, Distributed Computing, Developer Resource Hubs.

ما هي البدائل مفتوحة المصدر لـ oxnr/awesome-bigdata؟

تشمل البدائل مفتوحة المصدر لـ oxnr/awesome-bigdata: awesome-selfhosted/awesome-selfhosted — This project is a community-curated directory of open-source software designed for deployment in private server… cjbarber/toolsofthetrade — ToolsOfTheTrade is a comprehensive, curated directory designed to help developers and operational teams discover… vonng/ddia — This project serves as a comprehensive technical reference for the architecture and design of data-intensive… hazelcast/hazelcast — Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to… sindresorhus/awesome — This project is a community-maintained directory that serves as a comprehensive index of software tools, frameworks,… mahmoud/awesome-python-applications — This project is a curated directory and reference library of open-source Python applications. It serves as a…