awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目关于排名机制媒体报道MCP 服务器
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
oxnr avatar

oxnr/awesome-bigdata

0
View on GitHub↗
14,454 星标·2,582 分支·MIT·8 次浏览github.com/onurakpolat/awesome-bigdata↗

Awesome Bigdata

This project is a curated directory of software, frameworks, and educational resources designed for building, scaling, and maintaining distributed data processing and storage architectures. It serves as a comprehensive index for the distributed computing ecosystem, helping users identify the appropriate tools for managing large-scale information systems.

The repository functions as a central hub for data engineering, offering categorized access to technologies that support batch and stream processing, machine learning, and interactive querying. By organizing these resources, it assists in the design and development of complex data pipelines and the selection of infrastructure components for massive datasets.

Features

  • Curated Software Directories - Acts as a structured directory of high-quality software projects and frameworks organized by technical domain for distributed data architectures.
  • Data Analytics Engines - Provides a comprehensive index of high-performance computational engines for executing complex analytical queries on massive datasets.
  • Awesome List - A community-curated directory that catalogs and links out to other open-source projects, rather than a standalone tool you run yourself.
  • Large-Scale Training Frameworks - Curates infrastructure and orchestration frameworks for scaling machine learning model training across massive compute clusters.
  • Data Discovery Tools - Helps identify and evaluate databases and processing frameworks for large-scale data infrastructure.
  • Data Pipeline Orchestration - Schedules recurring batch jobs and resolves task dependencies for reliable data pipeline execution.
  • Distributed Computing - Executes data processing tasks across interconnected nodes to handle massive datasets through parallel computation.
  • Developer Resource Hubs - Serves as a centralized hub for accessing educational materials, libraries, and tools for designing large-scale data pipelines.
  • Columnar Storage Engines - Organizes data into columns to optimize analytical query performance and compression on massive datasets.
  • Distributed Computing Engines - Indexes a wide range of distributed computing engines and frameworks for batch, stream, and interactive data processing.
  • Stream Processing Systems - Processes continuous data flows through distributed buffers for real-time analysis and state updates.
  • Distributed Data Processing - Executes batch and real-time data workflows across computing clusters using parallel programming models.
  • High-Performance Data Infrastructures - Maintains scalable, high-performance storage systems for structured and unstructured data across cloud environments.
  • Log-Structured Storage - Uses append-only file structures to optimize write throughput and retrieval in distributed storage systems.
  • Query Languages - Facilitates interactive analysis of large datasets using standard query languages.
  • Engineering Guides - Offers educational resources and libraries for designing and maintaining distributed data processing architectures.
  • Directed Acyclic Graph Engines - Manages complex data pipeline dependencies and execution order across distributed environments.
  • Machine Learning Training - Supports training predictive models on large distributed datasets to improve analytical accuracy.
  • Big Data - Listed in the “Big Data” section of the Awesome awesome list.
  • Awesome Lists - Resources for big data frameworks and tools.
  • Data Ingestion - Collects and synchronizes streaming data from external sources into storage or processing systems.
  • Data Visualization Dashboards - Enables the creation of interactive dashboards and charts to visualize complex data insights.
  • Access Control - Implements authentication and authorization controls to protect sensitive data within distributed environments.
  • Role-Based Access Control - Enforces security by mapping user identities to specific permissions for data and infrastructure access.
  • Query Optimization Engines - Translates high-level analytical requests into optimized execution plans for distributed processing engines.

Star 历史

oxnr/awesome-bigdata 的 Star 历史图表oxnr/awesome-bigdata 的 Star 历史图表

AI 搜索

探索更多 awesome 仓库

用简单的语言描述您的需求 —— AI 将根据相关性为您从数千个精选开源项目中进行排序。

Start searching with AI

Awesome Bigdata 的开源替代方案

相似的开源项目,按与 Awesome Bigdata 的功能重合度排序。
  • awesome-selfhosted/awesome-selfhostedawesome-selfhosted 的头像

    awesome-selfhosted/awesome-selfhosted

    299,516在 GitHub 上查看↗

    This project is a community-curated directory of open-source software designed for deployment in private server environments and home labs. It serves as a comprehensive resource for discovering independent, self-hosted alternatives to mainstream cloud services, enabling users to maintain full data ownership and control over their digital infrastructure. The directory is structured through a hierarchical taxonomy that organizes a vast collection of applications into logical categories, ranging from media management and data analytics to private communication and team productivity tools. It dis

    awesomeawesome-listcloud
    在 GitHub 上查看↗299,516
  • cjbarber/toolsofthetradecjbarber 的头像

    cjbarber/ToolsOfTheTrade

    17,075在 GitHub 上查看↗

    ToolsOfTheTrade is a comprehensive, curated directory designed to help developers and operational teams discover software services for a wide range of technical and business requirements. It functions as a centralized resource catalog, indexing tools that support the entire software development lifecycle, infrastructure management, and organizational productivity. The project distinguishes itself by providing a structured index of both hosted and self-hosted solutions, enabling users to identify specific utilities for tasks ranging from continuous integration and project planning to applicati

    在 GitHub 上查看↗17,075
  • vonng/ddiaVonng 的头像

    Vonng/ddia

    22,648在 GitHub 上查看↗

    This project serves as a comprehensive technical reference for the architecture and design of data-intensive applications. It provides a structured analysis of the fundamental principles required to build reliable, scalable, and maintainable software systems, covering the core trade-offs inherent in modern data infrastructure. The repository explores the mechanics of distributed data management, including strategies for replication, partitioning, and achieving consensus across multiple nodes. It details the design of storage engines, indexing techniques, and transaction management models, whi

    Pythonbookdatabaseddia
    在 GitHub 上查看↗22,648
  • hazelcast/hazelcasthazelcast 的头像

    hazelcast/hazelcast

    6,570在 GitHub 上查看↗

    Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to support real-time analytics and event-driven applications. It functions as a partitioned, distributed key-value store that replicates data across cluster nodes to provide low-latency access and high availability. The platform also serves as a distributed SQL query engine, allowing users to execute standard SQL statements against both in-memory datasets and external data sources. What distinguishes Hazelcast is its use of a distributed consensus subsystem to maintain strongly consis

    Javabig-datacachingdata-in-motion
    在 GitHub 上查看↗6,570
查看 Awesome Bigdata 的所有 30 个替代方案→

常见问题解答

oxnr/awesome-bigdata 是做什么的?

This project is a curated directory of software, frameworks, and educational resources designed for building, scaling, and maintaining distributed data processing and storage architectures. It serves as a comprehensive index for the distributed computing ecosystem, helping users identify the appropriate tools for managing large-scale information systems.

oxnr/awesome-bigdata 的主要功能有哪些?

oxnr/awesome-bigdata 的主要功能包括:Curated Software Directories, Data Analytics Engines, Awesome List, Large-Scale Training Frameworks, Data Discovery Tools, Data Pipeline Orchestration, Distributed Computing, Developer Resource Hubs。

oxnr/awesome-bigdata 有哪些开源替代品?

oxnr/awesome-bigdata 的开源替代品包括: awesome-selfhosted/awesome-selfhosted — This project is a community-curated directory of open-source software designed for deployment in private server… cjbarber/toolsofthetrade — ToolsOfTheTrade is a comprehensive, curated directory designed to help developers and operational teams discover… vonng/ddia — This project serves as a comprehensive technical reference for the architecture and design of data-intensive… hazelcast/hazelcast — Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to… sindresorhus/awesome — This project is a community-maintained directory that serves as a comprehensive index of software tools, frameworks,… mahmoud/awesome-python-applications — This project is a curated directory and reference library of open-source Python applications. It serves as a…