awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

3 个仓库

Awesome GitHub RepositoriesData Parallelism Frameworks

Libraries that abstract the partitioning of data collections into independent units for concurrent execution.

Distinct from Parallel Work Partitioning: Distinct from general parallel work partitioning: focuses on the framework-level abstraction for data-parallel iterators.

Explore 3 awesome GitHub repositories matching devops & infrastructure · Data Parallelism Frameworks. Refine with filters or upvote what's useful.

Awesome Data Parallelism Frameworks GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • dask/daskdask 的头像

    dask/dask

    13,746在 GitHub 上查看↗

    Dask 是一个并行计算框架和分布式任务调度器,旨在将 Python 数据科学工作流从单机扩展到大型集群。它作为一个集群资源管理器,通过将任务及其依赖项表示为有向无环图来编排计算逻辑。这种架构允许系统在管理复杂执行要求的同时,自动将工作负载分配到可用硬件上。 该项目通过一个延迟评估引擎脱颖而出,该引擎将数据操作推迟到明确请求时才执行,从而实现全局图优化和高效的资源分配。它结合了内存感知数据溢出功能,以防止在处理超过可用内存的数据集时系统崩溃,并利用任务图融合将操作序列组合成单个执行步骤,从而最大限度地减少调度开销和节点间通信。 该平台为大规模数据分析提供了全面的功能面,包括对分布式机器学习、高性能计算集成和并行数据处理的支持。它提供了用于集群生命周期管理、性能分析和任务执行实时监控的广泛工具。用户可以在各种基础设施上部署这些环境,包括本地硬件、云提供商、容器化系统和高性能计算集群。

    Organizes large datasets into partitioned arrays and dataframes to enable parallel processing across distributed clusters.

    Pythondasknumpypandas
    在 GitHub 上查看↗13,746
  • rayon-rs/rayonrayon-rs 的头像

    rayon-rs/rayon

    13,071在 GitHub 上查看↗

    Rayon is a data parallelism library for Rust that provides a framework for converting sequential computations into parallel operations. It enables the transformation of standard data structures and loops into parallel iterators, allowing workloads to be distributed across multiple processor cores. By utilizing a work-stealing scheduler, the library dynamically balances tasks to maximize throughput and minimize execution time. The library distinguishes itself through its focus on safe, scoped task synchronization, which ensures that all spawned operations complete before a scope exits to preve

    Provides a framework for transforming sequential data structures into parallel iterators for concurrent processing.

    Rust
    在 GitHub 上查看↗13,071
  • sfu-db/connector-xsfu-db 的头像

    sfu-db/connector-x

    2,561在 GitHub 上查看↗

    Connector-X is a high-performance SQL data extraction library and bridge for transferring relational database records into memory-efficient data structures. It functions as a parallel database connector and federated query engine capable of executing and joining queries across multiple remote database connections to aggregate data locally. The project distinguishes itself through a zero-copy approach to data loading, which transfers SQL query results into memory structures without duplicating data. It maximizes throughput by partitioning SQL queries into threads, employing parallel columnar a

    Increases data throughput by splitting SQL queries into partitions and downloading them via multiple simultaneous threads.

    Rustcppdatabasedataframe
    在 GitHub 上查看↗2,561
  1. Home
  2. DevOps & Infrastructure
  3. Load Balancing
  4. Partitioning Algorithms
  5. Parallel Work Partitioning
  6. Data Parallelism Frameworks

探索子标签

  • Parallel SQL LoadingPartitioning SQL queries to download data across simultaneous threads for increased ingestion speed. **Distinct from Data Parallelism Frameworks:** Distinct from Data Parallelism Frameworks: focuses specifically on the partitioning and parallel downloading of SQL query results.