awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to waylau/apache-spark-tutorial

Open-source alternatives to Apache Spark Tutorial

20 open-source projects similar to waylau/apache-spark-tutorial, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Apache Spark Tutorial alternative.

  • akka/alpakka-kafkaakka avatar

    akka/alpakka-kafka

    1,421View on GitHub↗

    Alpakka Kafka connector - Alpakka is a Reactive Enterprise Integration library for Java and Scala, based on Reactive Streams and Akka.

    Scala
    View on GitHub↗1,421
  • ambster-public/awesome-qlikambster-public avatar

    ambster-public/awesome-qlik

    78View on GitHub↗

    A curated list of awesome Qlik extensions and resources for Qlik Sense and QlikView

    analyticsawesomeawesome-list
    View on GitHub↗78
  • apache/flinkapache avatar

    apache/flink

    26,086View on GitHub↗

    Apache Flink is a distributed processing engine designed for both high-throughput, low-latency data streams and finite batch workloads. It functions as a stateful stream processor and a SQL stream processing engine, providing a unified runtime to execute relational queries and event-based transformations. The system is distinguished by its ability to manage persistent operator state to ensure exactly-once processing guarantees and consistency during failures. It features specialized capabilities for complex event processing to detect temporal patterns and handles out-of-order events using eve

    Java
    View on GitHub↗26,086
  • apache/kafkaapache avatar

    apache/kafka

    32,846View on GitHub↗

    Kafka is a distributed event streaming platform designed for capturing, storing, and processing real-time data streams across interconnected nodes. It functions as a distributed commit log, providing a fault-tolerant storage mechanism that records state changes sequentially to ensure data consistency and durability across distributed environments. The platform distinguishes itself through a partitioned commit log architecture that enables horizontal scaling and parallel processing of data streams. It integrates a stream processing engine for continuous transformations and aggregations, while

    Javakafkascala
    View on GitHub↗32,846

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • apache/sparkapache avatar

    apache/spark

    43,467View on GitHub↗

    Apache Spark is a unified distributed data processing engine designed for large-scale data analysis and computation graphs. It functions as a distributed machine learning framework, a graph processing system, a real-time stream processor, and a SQL analytics engine. The system enables the execution of distributed SQL querying, large-scale graph analysis, and real-time stream analytics across clusters of machines. It also provides a scalable environment for implementing machine learning algorithms and predictive model development on massive datasets. The engine incorporates relational query e

    Scalabig-datajavajdbc
    View on GitHub↗43,467
  • awesome-spark/awesome-sparkawesome-spark avatar

    awesome-spark/awesome-spark

    1,882View on GitHub↗

    A curated list of awesome Apache Spark packages and resources.

    Shellapache-sparkawesomepyspark
    View on GitHub↗1,882
  • awesomedata/awesome-public-datasetsawesomedata avatar

    awesomedata/awesome-public-datasets

    75,979View on GitHub↗

    This project is a community-maintained, open-access directory of high-quality public datasets. It serves as a centralized reference point for researchers, developers, and data scientists to locate reliable information sources across a wide spectrum of industries and scientific fields. By providing a structured index, the repository facilitates the discovery of data necessary for exploratory analysis, machine learning model training, and the development of data-intensive applications. The directory distinguishes itself through a lightweight, platform-agnostic approach to resource indexing that

    aaron-swartzawesome-public-datasetsdatasets
    View on GitHub↗75,979
  • briatte/awesome-network-analysisbriatte avatar

    briatte/awesome-network-analysis

    4,067View on GitHub↗

    A curated list of awesome network analysis resources.

    Rawesomeawesome-listcomplex-networks
    View on GitHub↗4,067
  • datawhalechina/wonderful-sqlD

    datawhalechina/wonderful-sql

    0View on GitHub↗
    View on GitHub↗0
  • elasticsearch/elasticsearch-definitive-guideE

    elasticsearch/elasticsearch-definitive-guide

    0View on GitHub↗
    View on GitHub↗0
  • galliaproject/gallia-coregalliaproject avatar

    galliaproject/gallia-core

    88View on GitHub↗

    A schema-aware Scala library for data transformation

    Scala
    View on GitHub↗88
  • igorbarinov/awesome-data-engineeringigorbarinov avatar

    igorbarinov/awesome-data-engineering

    8,306View on GitHub↗
    awesomeawesome-list
    View on GitHub↗8,306
  • looly/elasticsearch-definitive-guide-cnlooly avatar

    looly/elasticsearch-definitive-guide-cn

    2,046View on GitHub↗

    Elasticsearch权威指南中文版

    Shell
    View on GitHub↗2,046
  • manuzhang/awesome-streamingmanuzhang avatar

    manuzhang/awesome-streaming

    2,986View on GitHub↗

    a curated list of awesome streaming frameworks, applications, etc

    awesomeawesome-listlist
    View on GitHub↗2,986
  • openmole/gridscaleopenmole avatar

    openmole/gridscale

    30View on GitHub↗

    Scala library for accessing various file, batch systems, job schedulers and grid middlewares.

    Scala
    View on GitHub↗30
  • oxnr/awesome-bigdataoxnr avatar

    oxnr/awesome-bigdata

    14,454View on GitHub↗

    This project is a curated directory of software, frameworks, and educational resources designed for building, scaling, and maintaining distributed data processing and storage architectures. It serves as a comprehensive index for the distributed computing ecosystem, helping users identify the appropriate tools for managing large-scale information systems. The repository functions as a central hub for data engineering, offering categorized access to technologies that support batch and stream processing, machine learning, and interactive querying. By organizing these resources, it assists in the

    awesomeawesome-listbigdata
    View on GitHub↗14,454
  • sduff/awesome-splunksduff avatar

    sduff/awesome-splunk

    159View on GitHub↗

    A collection of awesome resources for Splunk

    awesomeawesome-listsplunk
    View on GitHub↗159
  • spotify/sciospotify avatar

    spotify/scio

    2,628View on GitHub↗

    A Scala API for Apache Beam and Google Cloud Dataflow.

    Scala
    View on GitHub↗2,628
  • touk/nussknackerTouK avatar

    TouK/nussknacker

    731View on GitHub↗

    Low-code tool for automating actions on real time data | Stream processing for the users.

    Scala
    View on GitHub↗731
  • youngwookim/awesome-hadoopyoungwookim avatar

    youngwookim/awesome-hadoop

    1,117View on GitHub↗

    A curated list of amazingly awesome Hadoop and Hadoop ecosystem resources

    View on GitHub↗1,117