awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
apache avatar

apache/hadoop

0
View on GitHub↗
15,567 stars·9,221 forks·Java·Apache-2.0·28 viewshadoop.apache.org↗

Hadoop

Hadoop is a big data infrastructure suite and distributed data processing framework designed to store and process massive datasets across clusters of computers. It consists of a distributed storage system for managing large files across multiple nodes and a parallel computing engine for processing data across a distributed cluster.

The framework implements a distributed file system to ensure fault tolerance and high throughput, paired with a programming model that processes large datasets in parallel. It manages the underlying hardware and software environment required for distributed big data storage and cluster computing.

The system covers large scale data processing and big data infrastructure management. It provides capabilities for distributing data across clusters and executing computational tasks across multiple nodes to handle volumes of information too large for a single computer.

Features

  • Big Data Processing - Provides the primary infrastructure for managing, storing, and processing massive volumes of data across distributed systems.
  • Distributed Computing - Executes large-scale data analytics and processing tasks in parallel across distributed computing clusters.
  • Distributed Data Processing Frameworks - Provides a framework for partitioning, transforming, and processing large-scale datasets across distributed clusters.
  • Distributed File Systems - Implements a scalable distributed file system that splits files into blocks across commodity hardware nodes.
  • Large-Scale Data Computation - Implements a distributed framework for executing complex data analysis and computation across large clusters.
  • MapReduce Processing Engines - Provides a parallel computing engine based on the MapReduce programming model for processing massive datasets.
  • Distributed Storage Clusters - Creates a scalable storage system by aggregating multiple nodes into a unified distributed storage cluster.
  • Distributed File Systems - Implements a scalable distributed file system that manages large files across multiple nodes for fault tolerance.
  • Data-Locality Scheduling - Implements scheduling that minimizes network traffic by executing logic on the physical node where the required data resides.
  • Fault Tolerance - Ensures high data availability and resilience by replicating data blocks across multiple physical nodes.
  • Cluster Resource Managers - Provides a mechanism to allocate computing resources and schedule jobs across distributed network nodes to optimize hardware usage.
  • Master-Worker Coordination - Uses a central node to manage metadata and orchestrate task assignment to worker nodes across the cluster.
  • Big Data Frameworks - Framework for distributed processing of large datasets.
  • Data Processing - Distributed processing framework for big data workloads.
  • Data Processing and Analysis - Foundation for distributed storage and large-scale data processing.
  • Distributed Filesystems - Distributed filesystem for high-throughput application data.
  • Data Engineering - Framework for distributed processing of large datasets.
  • Data Infrastructure Management - Framework for distributed processing of large datasets across compute clusters.

Star history

Star history chart for apache/hadoopStar history chart for apache/hadoop

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Hadoop

These projects share indexed features with Hadoop. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • apache/flinkapache avatar

    apache/flink

    26,086View on GitHub↗

    Apache Flink is a distributed processing engine designed for both high-throughput, low-latency data streams and finite batch workloads. It functions as a stateful stream processor and a SQL stream processing engine, providing a unified runtime to execute relational queries and event-based transformations. The system is distinguished by its ability to manage persistent operator state to ensure exactly-once processing guarantees and consistency during failures. It features specialized capabilities for complex event processing to detect temporal patterns and handles out-of-order events using eve

    Java
    View on GitHub↗26,086
  • apache/hbaseapache avatar

    apache/hbase

    5,540View on GitHub↗

    HBase is a distributed, wide-column NoSQL store and big data storage engine designed for sparse datasets. It functions as a scalable columnar database built on top of the Hadoop Distributed File System to provide real-time read and write access to massive volumes of structured and unstructured data. The system acts as a cross-language database gateway, offering connectivity through native remote procedure calls, REST, and Thrift interfaces. It distinguishes itself through a master-worker coordination model that enables horizontal scaling and fault tolerance across a cluster. The project cove

    Java
    View on GitHub↗5,540
  • apache/sparkapache avatar

    apache/spark

    43,467View on GitHub↗

    Apache Spark is a unified distributed data processing engine designed for large-scale data analysis and computation graphs. It functions as a distributed machine learning framework, a graph processing system, a real-time stream processor, and a SQL analytics engine. The system enables the execution of distributed SQL querying, large-scale graph analysis, and real-time stream analytics across clusters of machines. It also provides a scalable environment for implementing machine learning algorithms and predictive model development on massive datasets. The engine incorporates relational query e

    Scalabig-datajavajdbc
    View on GitHub↗43,467
  • hazelcast/hazelcasthazelcast avatar

    hazelcast/hazelcast

    6,570View on GitHub↗

    Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to support real-time analytics and event-driven applications. It functions as a partitioned, distributed key-value store that replicates data across cluster nodes to provide low-latency access and high availability. The platform also serves as a distributed SQL query engine, allowing users to execute standard SQL statements against both in-memory datasets and external data sources. What distinguishes Hazelcast is its use of a distributed consensus subsystem to maintain strongly consis

    Javabig-datacachingdata-in-motion
    View on GitHub↗6,570
Compare all 30 related projects→

Frequently asked questions

What does apache/hadoop do?

Hadoop is a big data infrastructure suite and distributed data processing framework designed to store and process massive datasets across clusters of computers. It consists of a distributed storage system for managing large files across multiple nodes and a parallel computing engine for processing data across a distributed cluster.

What are the main features of apache/hadoop?

The main features of apache/hadoop are: Big Data Processing, Distributed Computing, Distributed Data Processing Frameworks, Distributed File Systems, Large-Scale Data Computation, MapReduce Processing Engines, Distributed Storage Clusters, Data-Locality Scheduling.

Which projects share features with apache/hadoop?

Projects with overlapping indexed features include: apache/flink — Apache Flink is a distributed processing engine designed for both high-throughput, low-latency data streams and finite… apache/hbase — HBase is a distributed, wide-column NoSQL store and big data storage engine designed for sparse datasets. It functions… apache/spark — Apache Spark is a unified distributed data processing engine designed for large-scale data analysis and computation… hazelcast/hazelcast — Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to… mahmoudparsian/data-algorithms-book — This repository is a collection of reference implementations and distributed data processing algorithms implemented in… jerrylead/sparkinternals — SparkInternals is a technical reference and architecture guide detailing the internal design and implementation of the…