awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
databricks avatar

databricks/learning-spark

0
View on GitHub↗
3,899 stars·2,416 forks·Java·mit·18 views

Learning Spark

This project is a learning curriculum and programming guide for Apache Spark, providing a structured set of educational resources and practical code examples for mastering distributed data processing. It serves as a course for building scalable data workflows and big data engineering pipelines.

The repository provides practical source code and project layouts that demonstrate how to connect external data stores, process streaming data, and organize code for distributed environments. It includes implementation examples for scaling machine learning algorithms across clusters to handle large training datasets.

The content covers the development of data workflows, the integration of external storage systems, and the process of compiling and packaging source code into executable assemblies for cluster deployment.

Features

  • How-To Structured Data - Provides practical code samples and functional examples demonstrating distributed data processing patterns.
  • Distributed Computing Curricula - Offers a structured educational course for mastering scalable data workflows and machine learning pipelines using Apache Spark.
  • Big Data Processing - Provides frameworks and methodologies for transforming massive volumes of data across distributed systems.
  • Data Processing Workflows - Guides the definition and execution of complex sequences of data analysis and transformation tasks.
  • Distributed Data Processing Frameworks - Implements systems for partitioning, transforming, and processing large-scale datasets across compute clusters.
  • Distributed Task Schedulers - Provides implementation patterns for orchestrating and distributing data processing workflows across computing clusters.
  • External Data Connectors - Demonstrates how to integrate and host external data streams using specific connectors for distributed processing.
  • External Storage Integrations - Implements support for connecting diverse external storage drivers to distributed processing engines.
  • Distributed Job Execution - Demonstrates how to execute computational jobs across multiple worker nodes using submission scripts.
  • Big Data Learning Paths - Provides a comprehensive set of educational resources and practical examples for mastering distributed data processing.
  • Code Examples - Offers practical source code and project layouts demonstrating distributed data and streaming processing.
  • Distributed Training - Provides implementation examples for scaling machine learning algorithms across clusters to handle massive training sets.
  • Scalable Distributed Pipelines - Demonstrates the development of high-scale data processing sequences across distributed compute resources.
  • Lazy Evaluation Frameworks - Illustrates the use of lazy evaluation frameworks to defer computation and enable global query optimization.
  • Machine Learning Pipelines - Implements scalable machine learning pipelines for distributed data transformation and model execution.
  • Orchestrator-Worker Models - Explains the architectural separation between central coordination logic and remote execution nodes in a cluster.
  • Polyglot Application Development - Shows how to implement processing functions across multiple languages through a shared core engine.

Star history

Star history chart for databricks/learning-sparkStar history chart for databricks/learning-spark

How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does databricks/learning-spark do?

This project is a learning curriculum and programming guide for Apache Spark, providing a structured set of educational resources and practical code examples for mastering distributed data processing. It serves as a course for building scalable data workflows and big data engineering pipelines.

What are the main features of databricks/learning-spark?

The main features of databricks/learning-spark are: How-To Structured Data, Distributed Computing Curricula, Big Data Processing, Data Processing Workflows, Distributed Data Processing Frameworks, Distributed Task Schedulers, External Data Connectors, External Storage Integrations.

What are some open-source alternatives to databricks/learning-spark?

Open-source alternatives to databricks/learning-spark include: mahmoudparsian/data-algorithms-book — This repository is a collection of reference implementations and distributed data processing algorithms implemented in… databricks/spark-the-definitive-guide — This project is an educational resource and technical manual for Apache Spark, focused on the architecture and… hazelcast/hazelcast — Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to… spotify/luigi — Luigi is a Python framework designed for building and managing complex batch data pipelines. It functions as a… zenml-io/zenml — ZenML is an orchestration platform designed for building, deploying, and monitoring reproducible machine learning… azkaban/azkaban — Azkaban is a distributed workflow manager and DAG-based job orchestrator designed as an enterprise batch processor. It…

Open-source alternatives to Learning Spark

Similar open-source projects, ranked by how many features they share with Learning Spark.
  • mahmoudparsian/data-algorithms-bookmahmoudparsian avatar

    mahmoudparsian/data-algorithms-book

    1,081View on GitHub↗

    This repository is a collection of reference implementations and distributed data processing algorithms implemented in Java and Scala for cluster computing frameworks. It provides computational recipes for solving complex data processing problems, including large-scale dataset joins, aggregations, and word count tasks. The implementations cover both MapReduce paradigms and Apache Spark integrations, enabling programmatic job submission and execution across distributed node infrastructures. The collection includes specialized utilities for statistical analysis and text processing, such as data

    Javaapache-hadoopapache-sparkdata-algorithms
    View on GitHub↗1,081
  • databricks/spark-the-definitive-guidedatabricks avatar

    databricks/Spark-The-Definitive-Guide

    3,099View on GitHub↗

    This project is an educational resource and technical manual for Apache Spark, focused on the architecture and practical application of large-scale data processing. It serves as a guide for big data engineering and distributed computing, covering the principles of parallel processing and fault-tolerant data distribution. The material provides instructional content on designing distributed ETL pipelines and implementing data analysis workflows. It includes tutorials for polyglot data processing, offering patterns and examples for using Python, Scala, and Java within a unified environment. The

    Scala
    View on GitHub↗3,099
  • hazelcast/hazelcasthazelcast avatar

    hazelcast/hazelcast

    6,570View on GitHub↗

    Hazelcast is a distributed data platform that combines an in-memory data grid with a stream processing engine to support real-time analytics and event-driven applications. It functions as a partitioned, distributed key-value store that replicates data across cluster nodes to provide low-latency access and high availability. The platform also serves as a distributed SQL query engine, allowing users to execute standard SQL statements against both in-memory datasets and external data sources. What distinguishes Hazelcast is its use of a distributed consensus subsystem to maintain strongly consis

    Javabig-datacachingdata-in-motion
    View on GitHub↗6,570
  • spotify/luigispotify avatar

    spotify/luigi

    18,676View on GitHub↗

    Luigi is a Python framework designed for building and managing complex batch data pipelines. It functions as a workflow orchestration engine that organizes tasks into directed acyclic graphs, ensuring that jobs execute in the correct logical order based on their dependencies. By utilizing a centralized scheduler, the system coordinates task execution across distributed environments, tracks global workflow state, and prevents redundant processing by verifying the existence of output targets before triggering any work. The project distinguishes itself through a robust state-tracking mechanism t

    Pythonhadoopluigiorchestration-framework
    View on GitHub↗18,676
See all 30 alternatives to Learning Spark→