awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
databricks avatar

databricks/koalas

0
View on GitHub↗
3,373 stars·371 forks·Python·Apache-2.0·9 views

Koalas

Koalas: pandas API on Apache Spark

Features

  • Simplification Tools - Provides a Pandas API on Apache Spark for big data.
  • Data Interfaces - Pandas DataFrame API on Spark.
  • Data Pipelines and Orchestration - Pandas API implementation for distributed processing on Spark.
  • Scientific Computing Libraries - Pandas API implementation for Apache Spark.

Star history

Star history chart for databricks/koalasStar history chart for databricks/koalas

How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to Koalas

Similar open-source projects, ranked by how many features they share with Koalas.
  • jitsucom/jitsujitsucom avatar

    jitsucom/jitsu

    4,782View on GitHub↗

    Jitsu is an open-source customer data platform designed to orchestrate event data pipelines. It captures, transforms, and routes behavioral data from web and server sources into data warehouses and analytics tools, providing a unified infrastructure for managing event streams. The platform distinguishes itself through its focus on self-hosted, containerized operations that grant users full control over their data security and privacy. It features a robust identity resolution engine that stitches disparate user identifiers into persistent profiles across sessions and devices, alongside program

    TypeScriptbigqueryclickhousedata-collection
    View on GitHub↗4,782
  • astronomer/dag-factoryastronomer avatar

    astronomer/dag-factory

    1,440View on GitHub↗

    Dag-factory is a framework for constructing and managing Apache Airflow data pipelines through declarative configuration files. By replacing manual procedural code with structured YAML definitions, it enables the programmatic generation of complex workflow structures, task dependencies, and execution schedules. The project distinguishes itself by mapping configuration keys directly to Python class constructors and operators, allowing for the dynamic instantiation of objects and custom logic. It supports hierarchical configuration inheritance to standardize settings across environments and pro

    Pythonairflowapache-airflowdags
    View on GitHub↗1,440
  • alluxio/alluxioAlluxio avatar

    Alluxio/alluxio

    7,202View on GitHub↗

    Alluxio is a virtual distributed file system and data orchestration layer that serves as a high-performance caching layer between cloud storage and compute clusters. It acts as a distributed data cache designed to accelerate data access for large-scale analytics and machine learning workloads. The system provides a unified interface that presents multiple heterogeneous storage backends as a single coherent namespace. This allows for the unification of diverse storage systems, enabling computation engines to access data from different providers without changing application code. The project c

    Java
    View on GitHub↗7,202
  • aporia-ai/mlnotifyaporia-ai avatar

    aporia-ai/mlnotify

    343View on GitHub↗

    🔔 No need to keep checking your training - just one import line and you'll know the second it's done.

    Vue
    View on GitHub↗343
See all 30 alternatives to Koalas→

Frequently asked questions

What does databricks/koalas do?

Koalas: pandas API on Apache Spark

What are the main features of databricks/koalas?

The main features of databricks/koalas are: Simplification Tools, Data Interfaces, Data Pipelines and Orchestration, Scientific Computing Libraries.

What are some open-source alternatives to databricks/koalas?

Open-source alternatives to databricks/koalas include: astronomer/dag-factory — Dag-factory is a framework for constructing and managing Apache Airflow data pipelines through declarative… jitsucom/jitsu — Jitsu is an open-source customer data platform designed to orchestrate event data pipelines. It captures, transforms,… alluxio/alluxio — Alluxio is a virtual distributed file system and data orchestration layer that serves as a high-performance caching… apple/turicreate — This project is an automated machine learning framework and toolkit designed for training and tuning custom models for… aporia-ai/mlnotify — 🔔 No need to keep checking your training - just one import line and you'll know the second it's done. airbytehq/airbyte — Airbyte is a data integration platform designed to synchronize information between diverse applications, databases,…