awesome-repositories.com
Blog
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoAcerca deCómo clasificamosPrensaServidor MCP
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Eventual-Inc avatar

Eventual-Inc/Daft

0
View on GitHub↗

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI
daft.ai↗

Daft

Daft is a distributed dataframe library and multimodal data processor designed to handle large-scale structured and unstructured data. It functions as a vectorized execution engine that processes tables alongside images, audio, and video, utilizing a unified schema to manage diverse data types.

The project distinguishes itself by combining distributed data engineering with large-scale AI inference. It provides an AI data pipeline for batch-optimizing model prompts and generating high-dimensional text embeddings, while utilizing zero-copy memory sharing to execute custom Python functions without processing overhead.

Its capabilities extend across cloud data lakehouse connectivity, supporting open table formats like Iceberg, Delta Lake, and Hudi. The engine employs lazy-evaluated execution plans and sampling-based schema inference to manage datasets that exceed single-node memory, scaling workloads from local cores to distributed Kubernetes clusters.

The system further includes a comprehensive suite for data transformation, covering columnar aggregation, window functions, and geospatial manipulation, as well as specialized tools for audio transcription and video frame extraction.

Features

  • Multimodal Processing - Handles structured tables alongside unstructured media like images and audio within a single unified processing framework.
  • Distributed Dataframes - Provides a distributed dataframe library for processing large-scale structured and unstructured data across local cores or Kubernetes clusters.
  • Batch Inference Pipelines - Provides a system for processing large-scale multimodal datasets through AI models via batch-optimized inference pipelines.
  • Inference Scaling - Runs batch prompts and generates embeddings across distributed GPU clusters to process massive datasets.
5,225 estrellas·401 forks·Rust·apache-2.0·8 vistas
  • Multimodal Processing Engines - Ships a multimodal data processor that handles tables alongside images, audio, and video using vectorized execution.
  • Column Value Aggregations - Calculates summary statistics like sums and averages across multiple columns for a single row.
  • Columnar Analytics - Utilizes vectorized columnar processing on contiguous memory blocks to maximize hardware utilization.
  • Data Ingestion - Creates dataframes by loading and parsing data from in-memory sources, files, and external integrations.
  • Distributed Computing - Executes data processing tasks across a cluster of machines to handle datasets that exceed the memory of a single node.
  • Python-Defined Transformations - Executes custom Python functions directly on data using zero-copy memory sharing for high-performance transformations.
  • Dataset Aggregations - Performs functional aggregation and summary statistics across large distributed datasets.
  • Distributed Data Engines - Executes complex transformations and aggregations on large datasets that exceed the memory of a single machine.
  • Grouped Aggregations - Groups data by specific keys to calculate aggregate statistics like mean and count.
  • Lazy Evaluation Frameworks - Implements a lazy-evaluated execution plan that defers data transformations until results are explicitly requested.
  • Multimodal Data Loading - Reads structured and unstructured data from cloud storage and AI repositories into a unified framework.
  • Cross-Language Zero-Copy Passings - Employs zero-copy memory sharing to pass data between the core engine and Python functions without overhead.
  • Multimodal Unified Schemas - Provides a unified schema that manages structured tables alongside images, audio, and video.
  • Schema Inference - Automatically determines dataset structure through sampling without loading the entire file.
  • Vectorized Execution Engines - Implements a vectorized execution engine that optimizes memory usage and CPU efficiency for high-performance data transformations.
  • User-Defined Data Functions - Allows the execution of custom user-defined logic directly on data stored within dataframes.
  • Distributed Task Orchestration - Distributes data processing tasks across multiple machines to handle datasets that exceed single-node memory.
  • Distributed Data Workload Scaling - Transitions processing from local execution to distributed clusters via orchestration platforms.
  • Data Transformation Pipelines - Defines lazy-evaluated plans of operations for manipulating and computing data through multi-stage workflows.
  • AI Tool Execution - Executes model prompts and generates embeddings through optimized connections to external AI providers.
  • Audio Transcription - Converts audio files into textual segments with timestamps using speech-to-text models.
  • Synthetic Media Generators - Generates synthetic images from textual prompts using local GPU-accelerated diffusion models.
  • AI Model Integrations - Provides interfaces for connecting multimodal data processing pipelines to various local and cloud-based AI models.
  • Model Inference - Implements utilities for running model protocols, including text embedding, across multiple AI providers.
  • Text Embedding Generators - Generates high-dimensional vector representations of text using GPU acceleration for vector database storage.
  • Embedding Generation Pipelines - Generates high-dimensional text embeddings and calculates vector similarity for storage in vector search engines.
  • Data Deduplication Tools - Removes duplicate content from large text corpora using hashing algorithms.
  • Data Persistence and Storage - Persists processed datasets to local or remote destinations including Parquet and S3.
  • Cloud Data Lake Integrations - Provides connectivity for reading and writing data using open table formats like Iceberg and Delta Lake.
  • Data Partitioning - Divides large datasets into smaller segments using time-based or hash-based partitioning.
  • Data Processing - Provides general utilities for type casting, null filling, and conditional case-when expressions.
  • Data Source Connectivity Tools - Provides universal connectivity to data stored across cloud storage, table formats, and AI repositories.
  • Multi-Source Data Integration - Accesses data from diverse sources including cloud storage and enterprise catalogs without manual configuration.
  • Structured Types - Constructs structured data types from expressions and flattens nested fields into separate columns.
  • Date and Time Libraries - Performs temporal arithmetic and timezone conversions on timestamps.
  • Lakehouse Table Formats - Reads and writes data using open table formats such as Iceberg, Delta Lake, and Hudi.
  • Lazy Query Execution - Defines data transformations and schemas using lazy evaluation to optimize the processing pipeline.
  • List Processing Tools - Provides operations to filter, sort, flatten, and map elements within list columns.
  • Numeric Calculators - Provides a suite of mathematical functions including trigonometry, logarithms, and rounding.
  • Remote Query Execution - Distributes processing tasks across a remote compute cluster to leverage external hardware resources.
  • Window Functions - Implements context-aware window functions for complex calculations across sets of related rows.
  • Inference Batching - Parallelizes model prompts across local processor cores to maximize throughput for large multimodal datasets.
  • Inference Capabilities - Enables text and image classification and embedding generation via external model providers.
  • Kubernetes Deployments - Runs data processing scripts on Kubernetes clusters using both single-node and distributed setups.
  • Kubernetes Job Orchestration - Deploys and scales data processing jobs on Kubernetes clusters to leverage remote compute.
  • Cloud Storage Connectors - Interfaces with external storage providers and databases including S3 and various table formats.
  • Audio Processing - Extracts metadata and resamples audio files as part of a multimodal data processing pipeline.
  • Image Processing - Decodes images and extracts metadata to generate perceptual hashes for duplicate detection.
  • Video File Processors - Extracts metadata and captures specific frames from video files for analysis.
  • Row Windowing - Computes values across related rows to analyze local data trends within the dataframe.
  • Resource Orchestration - Prevents out-of-memory errors using vectorized execution and intelligent resource management.
  • Data Processing Libraries - Distributed dataframe engine for large-scale data.
  • Historial de estrellas

    Gráfico del historial de estrellas de eventual-inc/daftGráfico del historial de estrellas de eventual-inc/daft

    Preguntas frecuentes

    ¿Qué hace eventual-inc/daft?

    Daft is a distributed dataframe library and multimodal data processor designed to handle large-scale structured and unstructured data. It functions as a vectorized execution engine that processes tables alongside images, audio, and video, utilizing a unified schema to manage diverse data types.

    ¿Cuáles son las características principales de eventual-inc/daft?

    Las características principales de eventual-inc/daft son: Multimodal Processing, Distributed Dataframes, Batch Inference Pipelines, Inference Scaling, Multimodal Processing Engines, Column Value Aggregations, Columnar Analytics, Data Ingestion.

    ¿Qué alternativas de código abierto existen para eventual-inc/daft?

    Las alternativas de código abierto para eventual-inc/daft incluyen: dask/dask — Dask is a parallel computing framework and distributed task scheduler designed to scale Python data science workflows… prestodb/presto — Presto is a distributed SQL query engine designed for high-performance analytical processing across heterogeneous data… databricks/spark-the-definitive-guide — This project is an educational resource and technical manual for Apache Spark, focused on the architecture and… ray-project/ray — Ray is a distributed computing framework designed to scale Python and Java applications across clusters by abstracting… lancedb/lancedb — LanceDB is a vector database and columnar data store designed to function as a versioned dataset manager and vector… rapidsai/cudf — cuDF is a GPU-accelerated dataframe library and data processing engine designed for manipulating and analyzing large…

    Alternativas open-source a Daft

    Proyectos open-source similares, clasificados según cuántas características comparten con Daft.
    • dask/daskAvatar de dask

      dask/dask

      13,746Ver en GitHub↗

      Dask is a parallel computing framework and distributed task scheduler designed to scale Python data science workflows from single machines to large clusters. It functions as a cluster resource manager that orchestrates computational logic by representing tasks and their dependencies as directed acyclic graphs. This architecture allows the system to automate the distribution of workloads across available hardware while managing complex execution requirements. The project distinguishes itself through a lazy evaluation engine that defers data operations until they are explicitly requested, enabl

      Pythondasknumpypandas
      Ver en GitHub↗13,746
    • prestodb/prestoAvatar de prestodb

      prestodb/presto

      16,711Ver en GitHub↗

      Presto is a distributed SQL query engine designed for high-performance analytical processing across heterogeneous data sources. It functions as a data federation platform and massively parallel processing engine, allowing users to execute interactive queries against diverse storage systems without requiring data migration. By mapping remote metadata and structures to a unified relational namespace, it enables seamless cross-platform analysis through a standard SQL interface. The engine distinguishes itself through a pluggable connector architecture and a shared-nothing distributed processing

      Javabig-datadatahadoop
      Ver en GitHub↗16,711
    • databricks/spark-the-definitive-guideAvatar de databricks

      databricks/Spark-The-Definitive-Guide

      3,099Ver en GitHub↗

      This project is an educational resource and technical manual for Apache Spark, focused on the architecture and practical application of large-scale data processing. It serves as a guide for big data engineering and distributed computing, covering the principles of parallel processing and fault-tolerant data distribution. The material provides instructional content on designing distributed ETL pipelines and implementing data analysis workflows. It includes tutorials for polyglot data processing, offering patterns and examples for using Python, Scala, and Java within a unified environment. The

      Scala
      Ver en GitHub↗3,099
    • ray-project/rayAvatar de ray-project

      ray-project/ray

      42,895Ver en GitHub↗

      Ray is a distributed computing framework designed to scale Python and Java applications across clusters by abstracting task scheduling and resource management. It functions as a resource-aware execution engine that manages task dependencies, placement, and fault tolerance across networked compute nodes. At its core, the system provides a stateful actor model, allowing developers to define classes that run in dedicated processes to maintain and mutate internal state across remote method calls. The framework distinguishes itself through a robust cross-language interoperability layer, enabling f

      Pythondata-sciencedeep-learningdeployment
      Ver en GitHub↗42,895
    Ver las 30 alternativas a Daft→