awesome-repositories.com
Blog
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoAcerca deCómo clasificamosPrensaServidor MCP
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
databricks avatar

databricks/koalas

0
View on GitHub↗
3,373 estrellas·371 forks·Python·Apache-2.0·1 vista

Koalas

Koalas: pandas API on Apache Spark

Features

  • Simplification Tools - Provides a Pandas API on Apache Spark for big data.
  • Data Interfaces - Pandas DataFrame API on Spark.
  • Data Pipelines and Orchestration - Pandas API implementation for distributed processing on Spark.
  • Scientific Computing Libraries - Pandas API implementation for Apache Spark.

Historial de estrellas

Gráfico del historial de estrellas de databricks/koalasGráfico del historial de estrellas de databricks/koalas

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Alternativas open-source a Koalas

Proyectos open-source similares, clasificados según cuántas características comparten con Koalas.
  • jitsucom/jitsuAvatar de jitsucom

    jitsucom/jitsu

    4,782Ver en GitHub↗

    Jitsu is an open-source customer data platform designed to orchestrate event data pipelines. It captures, transforms, and routes behavioral data from web and server sources into data warehouses and analytics tools, providing a unified infrastructure for managing event streams. The platform distinguishes itself through its focus on self-hosted, containerized operations that grant users full control over their data security and privacy. It features a robust identity resolution engine that stitches disparate user identifiers into persistent profiles across sessions and devices, alongside program

    TypeScriptbigqueryclickhousedata-collection
    Ver en GitHub↗4,782
  • astronomer/dag-factoryAvatar de astronomer

    astronomer/dag-factory

    1,440Ver en GitHub↗

    Dag-factory is a framework for constructing and managing Apache Airflow data pipelines through declarative configuration files. By replacing manual procedural code with structured YAML definitions, it enables the programmatic generation of complex workflow structures, task dependencies, and execution schedules. The project distinguishes itself by mapping configuration keys directly to Python class constructors and operators, allowing for the dynamic instantiation of objects and custom logic. It supports hierarchical configuration inheritance to standardize settings across environments and pro

    Pythonairflowapache-airflowdags
    Ver en GitHub↗1,440
  • alluxio/alluxioAvatar de Alluxio

    Alluxio/alluxio

    7,202Ver en GitHub↗

    Alluxio is a virtual distributed file system and data orchestration layer that serves as a high-performance caching layer between cloud storage and compute clusters. It acts as a distributed data cache designed to accelerate data access for large-scale analytics and machine learning workloads. The system provides a unified interface that presents multiple heterogeneous storage backends as a single coherent namespace. This allows for the unification of diverse storage systems, enabling computation engines to access data from different providers without changing application code. The project c

    Java
    Ver en GitHub↗7,202
  • aporia-ai/mlnotifyAvatar de aporia-ai

    aporia-ai/mlnotify

    343Ver en GitHub↗

    🔔 No need to keep checking your training - just one import line and you'll know the second it's done.

    Vue
    Ver en GitHub↗343
Ver las 30 alternativas a Koalas→

Preguntas frecuentes

¿Qué hace databricks/koalas?

Koalas: pandas API on Apache Spark

¿Cuáles son las características principales de databricks/koalas?

Las características principales de databricks/koalas son: Simplification Tools, Data Interfaces, Data Pipelines and Orchestration, Scientific Computing Libraries.

¿Qué alternativas de código abierto existen para databricks/koalas?

Las alternativas de código abierto para databricks/koalas incluyen: astronomer/dag-factory — Dag-factory is a framework for constructing and managing Apache Airflow data pipelines through declarative… jitsucom/jitsu — Jitsu is an open-source customer data platform designed to orchestrate event data pipelines. It captures, transforms,… alluxio/alluxio — Alluxio is a virtual distributed file system and data orchestration layer that serves as a high-performance caching… apple/turicreate — This project is an automated machine learning framework and toolkit designed for training and tuning custom models for… aporia-ai/mlnotify — 🔔 No need to keep checking your training - just one import line and you'll know the second it's done. airbytehq/airbyte — Airbyte is a data integration platform designed to synchronize information between diverse applications, databases,…