10 个仓库
Frameworks for processing large-scale datasets and distributed execution.
Explore 10 awesome GitHub repositories matching part of an awesome list · Big Data and Distributed Computing. Refine with filters or upvote what's useful.
Ray is a distributed computing framework designed to scale Python and Java applications across clusters by abstracting task scheduling and resource management. It functions as a resource-aware execution engine that manages task dependencies, placement, and fault tolerance across networked compute nodes. At its core, the system provides a stateful actor model, allowing developers to define classes that run in dedicated processes to maintain and mutate internal state across remote method calls. The framework distinguishes itself through a robust cross-language interoperability layer, enabling f
Distributed execution framework for scaling applications.
Dask 是一个并行计算框架和分布式任务调度器,旨在将 Python 数据科学工作流从单机扩展到大型集群。它作为一个集群资源管理器,通过将任务及其依赖项表示为有向无环图来编排计算逻辑。这种架构允许系统在管理复杂执行要求的同时,自动将工作负载分配到可用硬件上。 该项目通过一个延迟评估引擎脱颖而出,该引擎将数据操作推迟到明确请求时才执行,从而实现全局图优化和高效的资源分配。它结合了内存感知数据溢出功能,以防止在处理超过可用内存的数据集时系统崩溃,并利用任务图融合将操作序列组合成单个执行步骤,从而最大限度地减少调度开销和节点间通信。 该平台为大规模数据分析提供了全面的功能面,包括对分布式机器学习、高性能计算集成和并行数据处理的支持。它提供了用于集群生命周期管理、性能分析和任务执行实时监控的广泛工具。用户可以在各种基础设施上部署这些环境,包括本地硬件、云提供商、容器化系统和高性能计算集群。
Distributed computing and parallel dataframe processing.
CuPy 是一个 CUDA 数组计算库,实现了与 NumPy 兼容的接口,用于在 NVIDIA GPU 上执行数组操作和数值计算。它作为一个 GPU 加速数值库和基于 CUDA 的 SciPy 实现,将繁重的计算卸载到图形硬件上,以提高科学和工程工作负载的处理速度。 该库支持多框架张量交换,允许使用标准化的内存布局在不同的深度学习框架之间共享数据缓冲区,从而避免内存拷贝。它还支持自定义 GPU 内核集成,允许将数组数据连接到低级 API,以便精确控制硬件执行。 该项目广泛涵盖了高性能数组处理和科学计算工作流。其功能包括加速数组计算和提供大规模数值计算工具。
CUDA-accelerated NumPy-compatible array operations.
cuDF is a GPU-accelerated dataframe library and data processing engine designed for manipulating and analyzing large tabular datasets. It provides a high-level API for executing filtering, joining, and aggregating operations directly on GPU hardware. The project integrates the Apache Arrow memory format to enable zero-copy data transfers and includes a just-in-time compiler for executing custom user-defined functions on the GPU. The library features specialized acceleration for existing workflows by redirecting standard Pandas dataframe calls and Polars query plans to a GPU backend. It also p
GPU-accelerated dataframe library.
h2o-3 is a distributed machine learning platform and automated machine learning framework designed for training and deploying predictive models using distributed in-memory computing. It functions as a deep learning framework and a distributed model scoring engine, capable of operating as a Kubernetes ML cluster to process large datasets in parallel. The platform distinguishes itself through automated machine learning capabilities that automatically select the best algorithms and hyperparameters to optimize model performance. It provides specialized deep learning toolkits for tasks including i
Distributed machine learning and out-of-memory dataframes.
An implementation of chunked, compressed, N-dimensional arrays for Python.
Distributed storage for multi-dimensional arrays.
Petastorm library enables single machine or distributed training and evaluation of deep learning models from datasets in Apache Parquet format. It supports ML frameworks such as Tensorflow, Pytorch, and PySpark and can be used from pure Python code.
Data access library for parquet files.
Library for reading and writing large multi-dimensional arrays.
Reading and writing large multi-dimensional arrays.
Fast NumPy array functions written in C
Fast NumPy array functions implemented in C.