awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to spotify/scio

Open-source alternatives to Scio

29 open-source projects similar to spotify/scio, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Scio alternative.

  • apache/icebergapache avatar

    apache/iceberg

    8,972View on GitHub↗

    Iceberg is an open table format and big data table manager designed for huge analytic datasets in cloud storage. It provides a specification for tracking large-scale datasets to maintain transactional consistency and structural integrity. The project utilizes a standardized REST catalog interface to manage table metadata, ensuring interoperability between different compute engines. This allows diverse query engines to connect to a single table interface and maintain consistency across different processing frameworks. Its core capabilities include managing large-scale analytic tables, coordin

    Java
    View on GitHub↗8,972
  • apache/kylinapache avatar

    apache/kylin

    3,765View on GitHub↗

    Kylin is a distributed OLAP engine designed for executing fast SQL queries on massive datasets. It utilizes multi-dimensional data cubes to pre-calculate data aggregates, enabling sub-second response times for large-scale analytical queries and big data analytics. The system focuses on large-scale data warehousing and multi-dimensional data modeling. It allows for the organization and querying of vast amounts of structured data to support business intelligence and reporting workflows through distributed SQL querying.

    Javakylin
    View on GitHub↗3,765
  • stellar/stellar-corestellar avatar

    stellar/stellar-core

    3,269View on GitHub↗

    Stellar Core is the primary software implementation of the Stellar blockchain network, serving as a distributed ledger and a Federated Byzantine Agreement system. It functions as a core node that maintains the shared state of the network and provides a runtime environment for executing WebAssembly smart contracts. The project enables the creation and management of digital assets, including the implementation of decentralized exchanges through distributed orderbooks and automated liquidity pools. It facilitates cross-border payment settlement by routing assets via path payments and bridging di

    C++
    View on GitHub↗3,269

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • mahmoudparsian/data-algorithms-bookmahmoudparsian avatar

    mahmoudparsian/data-algorithms-book

    1,081View on GitHub↗

    This repository is a collection of reference implementations and distributed data processing algorithms implemented in Java and Scala for cluster computing frameworks. It provides computational recipes for solving complex data processing problems, including large-scale dataset joins, aggregations, and word count tasks. The implementations cover both MapReduce paradigms and Apache Spark integrations, enabling programmatic job submission and execution across distributed node infrastructures. The collection includes specialized utilities for statistical analysis and text processing, such as data

    Javaapache-hadoopapache-sparkdata-algorithms
    View on GitHub↗1,081
  • apache/sparkapache avatar

    apache/spark

    43,467View on GitHub↗

    Apache Spark is a unified distributed data processing engine designed for large-scale data analysis and computation graphs. It functions as a distributed machine learning framework, a graph processing system, a real-time stream processor, and a SQL analytics engine. The system enables the execution of distributed SQL querying, large-scale graph analysis, and real-time stream analytics across clusters of machines. It also provides a scalable environment for implementing machine learning algorithms and predictive model development on massive datasets. The engine incorporates relational query e

    Scalabig-datajavajdbc
    View on GitHub↗43,467
  • awesome-spark/awesome-sparkawesome-spark avatar

    awesome-spark/awesome-spark

    1,882View on GitHub↗

    A curated list of awesome Apache Spark packages and resources.

    Shellapache-sparkawesomepyspark
    View on GitHub↗1,882
  • awesomedata/awesome-public-datasetsawesomedata avatar

    awesomedata/awesome-public-datasets

    75,979View on GitHub↗

    This project is a community-maintained, open-access directory of high-quality public datasets. It serves as a centralized reference point for researchers, developers, and data scientists to locate reliable information sources across a wide spectrum of industries and scientific fields. By providing a structured index, the repository facilitates the discovery of data necessary for exploratory analysis, machine learning model training, and the development of data-intensive applications. The directory distinguishes itself through a lightweight, platform-agnostic approach to resource indexing that

    aaron-swartzawesome-public-datasetsdatasets
    View on GitHub↗75,979
  • briatte/awesome-network-analysisbriatte avatar

    briatte/awesome-network-analysis

    4,067View on GitHub↗

    A curated list of awesome network analysis resources.

    Rawesomeawesome-listcomplex-networks
    View on GitHub↗4,067
  • datawhalechina/wonderful-sqlD

    datawhalechina/wonderful-sql

    0View on GitHub↗
    View on GitHub↗0
  • elasticsearch/elasticsearch-definitive-guideE

    elasticsearch/elasticsearch-definitive-guide

    0View on GitHub↗
    View on GitHub↗0
  • galliaproject/gallia-coregalliaproject avatar

    galliaproject/gallia-core

    88View on GitHub↗

    A schema-aware Scala library for data transformation

    Scala
    View on GitHub↗88
  • googlecloudplatform/bigquery-utilsGoogleCloudPlatform avatar

    GoogleCloudPlatform/bigquery-utils

    1,303View on GitHub↗

    Useful scripts, udfs, views, and other utilities for migration and data warehouse operations in BigQuery.

    Jupyter Notebook
    View on GitHub↗1,303
  • googlecloudplatform/dataflowtemplatesGoogleCloudPlatform avatar

    GoogleCloudPlatform/DataflowTemplates

    1,296View on GitHub↗

    Cloud Dataflow Google-provided templates for solving in-Cloud data tasks

    Javaapache-beambigquerybigtable
    View on GitHub↗1,296
  • googlecloudplatform/psqGoogleCloudPlatform avatar

    GoogleCloudPlatform/psq

    211View on GitHub↗

    psq - Cloud Pub/Sub Task Queue for Python.

    Python
    View on GitHub↗211
  • igorbarinov/awesome-data-engineeringigorbarinov avatar

    igorbarinov/awesome-data-engineering

    8,306View on GitHub↗
    awesomeawesome-list
    View on GitHub↗8,306
  • looly/elasticsearch-definitive-guide-cnlooly avatar

    looly/elasticsearch-definitive-guide-cn

    2,046View on GitHub↗

    Elasticsearch权威指南中文版

    Shell
    View on GitHub↗2,046
  • manuzhang/awesome-streamingmanuzhang avatar

    manuzhang/awesome-streaming

    2,986View on GitHub↗

    a curated list of awesome streaming frameworks, applications, etc

    awesomeawesome-listlist
    View on GitHub↗2,986
  • openmole/gridscaleopenmole avatar

    openmole/gridscale

    30View on GitHub↗

    Scala library for accessing various file, batch systems, job schedulers and grid middlewares.

    Scala
    View on GitHub↗30
  • oxnr/awesome-bigdataoxnr avatar

    oxnr/awesome-bigdata

    14,454View on GitHub↗

    This project is a curated directory of software, frameworks, and educational resources designed for building, scaling, and maintaining distributed data processing and storage architectures. It serves as a comprehensive index for the distributed computing ecosystem, helping users identify the appropriate tools for managing large-scale information systems. The repository functions as a central hub for data engineering, offering categorized access to technologies that support batch and stream processing, machine learning, and interactive querying. By organizing these resources, it assists in the

    awesomeawesome-listbigdata
    View on GitHub↗14,454
  • sduff/awesome-splunksduff avatar

    sduff/awesome-splunk

    159View on GitHub↗

    A collection of awesome resources for Splunk

    awesomeawesome-listsplunk
    View on GitHub↗159
  • spotify/heroicspotify avatar

    spotify/heroic

    846View on GitHub↗

    The Heroic Time Series Database

    Java
    View on GitHub↗846
  • spotify/spark-bigqueryspotify avatar

    spotify/spark-bigquery

    156View on GitHub↗

    spark-bigquery

    Scala
    View on GitHub↗156
  • touk/nussknackerTouK avatar

    TouK/nussknacker

    731View on GitHub↗

    Low-code tool for automating actions on real time data | Stream processing for the users.

    Scala
    View on GitHub↗731
  • waylau/apache-spark-tutorialW

    waylau/apache-spark-tutorial

    0View on GitHub↗
    View on GitHub↗0
  • akka/alpakka-kafkaakka avatar

    akka/alpakka-kafka

    1,421View on GitHub↗

    Alpakka Kafka connector - Alpakka is a Reactive Enterprise Integration library for Java and Scala, based on Reactive Streams and Akka.

    Scala
    View on GitHub↗1,421
  • youngwookim/awesome-hadoopyoungwookim avatar

    youngwookim/awesome-hadoop

    1,117View on GitHub↗

    A curated list of amazingly awesome Hadoop and Hadoop ecosystem resources

    View on GitHub↗1,117
  • ambster-public/awesome-qlikambster-public avatar

    ambster-public/awesome-qlik

    78View on GitHub↗

    A curated list of awesome Qlik extensions and resources for Qlik Sense and QlikView

    analyticsawesomeawesome-list
    View on GitHub↗78
  • apache/flinkapache avatar

    apache/flink

    26,086View on GitHub↗

    Apache Flink is a distributed processing engine designed for both high-throughput, low-latency data streams and finite batch workloads. It functions as a stateful stream processor and a SQL stream processing engine, providing a unified runtime to execute relational queries and event-based transformations. The system is distinguished by its ability to manage persistent operator state to ensure exactly-once processing guarantees and consistency during failures. It features specialized capabilities for complex event processing to detect temporal patterns and handles out-of-order events using eve

    Java
    View on GitHub↗26,086
  • apache/kafkaapache avatar

    apache/kafka

    32,846View on GitHub↗

    Kafka is a distributed event streaming platform designed for capturing, storing, and processing real-time data streams across interconnected nodes. It functions as a distributed commit log, providing a fault-tolerant storage mechanism that records state changes sequentially to ensure data consistency and durability across distributed environments. The platform distinguishes itself through a partitioned commit log architecture that enables horizontal scaling and parallel processing of data streams. It integrates a stream processing engine for continuous transformations and aggregations, while

    Javakafkascala
    View on GitHub↗32,846