awesome-repositories.com
Blog
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetÀ proposNotre méthodologiePresseServeur MCP
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to nv-legate/legate.numpy

Open-source alternatives to Legate.numpy

28 open-source projects similar to nv-legate/legate.numpy, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Legate.numpy alternative.

  • actor-framework/actor-frameworkAvatar de actor-framework

    actor-framework/actor-framework

    3,426Voir sur GitHub↗

    An Open Source Implementation of the Actor Model in C++

    C++actor-modelactorsasync
    Voir sur GitHub↗3,426
  • alpaka-group/alpakaAvatar de alpaka-group

    alpaka-group/alpaka

    416Voir sur GitHub↗

    alpaka - Abstraction Library for Parallel Kernel Acceleration

    C++
    Voir sur GitHub↗416
  • bloomen/transwarpAvatar de bloomen

    bloomen/transwarp

    633Voir sur GitHub↗

    A header-only C++ library for task concurrency

    C++
    Voir sur GitHub↗633
  • datenlord/async-rdmaAvatar de datenlord

    datenlord/async-rdma

    438Voir sur GitHub↗

    A framework for writing RDMA applications with high-level abstraction and asynchronous APIs.

    Rust
    Voir sur GitHub↗438
  • exaloop/codonAvatar de exaloop

    exaloop/codon

    16,803Voir sur GitHub↗

    Codon is an LLVM-based Python compiler and statically typed implementation that translates source code into optimized machine instructions. It functions as a high-performance numerical backend and a GPU computing framework designed to remove runtime overhead. The project implements a compiled alternative to NumPy, translating array logic directly into machine code. It differentiates itself by generating specialized hardware kernels for graphics processors and utilizing static type inference to enable aggressive machine-code optimization. The system provides capabilities for parallel workload

    Python
    Voir sur GitHub↗16,803

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Find more with AI search
  • facebookincubator/dispensoAvatar de facebookincubator

    facebookincubator/dispenso

    282Voir sur GitHub↗

    The project provides high-performance concurrency, enabling highly parallel computation.

    C++
    Voir sur GitHub↗282
  • fastflow/fastflowAvatar de fastflow

    fastflow/fastflow

    311Voir sur GitHub↗

    FastFlow is a programming library implemented in modern C++ and targeting multi/many-cores (there exists an experimental version based on ZeroMQ targeting distributed systems). It offers both a set of high-level ready-to-use parallel patterns and a set of mechanisms and composable components…

    C++
    Voir sur GitHub↗311
  • google/highwayAvatar de google

    google/highway

    5,644Voir sur GitHub↗

    Highway is a portable C++ library and hardware abstraction layer designed for writing single instruction multiple data (SIMD) code. It provides a unified interface that maps data-parallel logic to various CPU instruction sets, enabling the development of high-performance software that runs across different processor architectures without requiring architecture-specific assembly. The project features a dynamic instruction dispatcher that selects the most efficient CPU instruction set at runtime based on detected hardware. It also supports static target specialization and extensible mechanisms

    C++
    Voir sur GitHub↗5,644
  • heteroflow/heteroflowAvatar de Heteroflow

    Heteroflow/Heteroflow

    109Voir sur GitHub↗

    A header-only C++ library to help you quickly write concurrent CPU-GPU programs using task models

    C++
    Voir sur GitHub↗109
  • horovod/horovodAvatar de horovod

    horovod/horovod

    14,686Voir sur GitHub↗

    Horovod is a distributed deep learning framework and gradient synchronizer designed to scale model training across multiple GPUs and compute nodes. It functions as a distributed training orchestrator and an elastic training engine, utilizing an MPI collective communication library to synchronize weights and gradients across TensorFlow, PyTorch, Keras, and MXNet models. The system distinguishes itself through dynamic elastic scaling, which allows it to adjust the number of active workers at runtime and recover from node failures. It optimizes communication efficiency using tensor fusion batchi

    Python
    Voir sur GitHub↗14,686
  • intelligentsoftwaresystems/galoisAvatar de IntelligentSoftwareSystems

    IntelligentSoftwareSystems/Galois

    353Voir sur GitHub↗

    Overview

    C++
    Voir sur GitHub↗353
  • ispc/ispcAvatar de ispc

    ispc/ispc

    2,843Voir sur GitHub↗

    ISPC is a vectorizing compiler and SIMD parallel programming language that implements a single program multiple data model. It serves as a toolchain for translating C-based code with parallel extensions into optimized machine code for various CPU and GPU architectures using an LLVM backend. The compiler is designed for cross-platform SIMD toolchain support, generating specialized instruction sets for x86 SSE/AVX, ARM NEON, and Intel GPU from a single source. It features a runtime dispatch mechanism that selects the most efficient hardware-specific implementation for the current system during

    C++compilerintelispc
    Voir sur GitHub↗2,843
  • it4innovations/hyperqueueI

    It4innovations/hyperqueue

    0Voir sur GitHub↗

    HyperQueue is a tool designed to simplify execution of large workflows (task graphs) on HPC clusters. It allows you to execute a large number of tasks in a simple way, without having to manually submit jobs into batch schedulers like Slurm or PBS. You specify what you want to compute and…

    Voir sur GitHub↗0
  • kokkos/kokkosAvatar de kokkos

    kokkos/kokkos

    2,571Voir sur GitHub↗

    Kokkos C++ Performance Portability Programming Ecosystem: The Programming Model - Parallel Execution and Memory Abstraction

    C++
    Voir sur GitHub↗2,571
  • komputeproject/komputeAvatar de KomputeProject

    KomputeProject/kompute

    2,519Voir sur GitHub↗

    General purpose GPU compute framework built on Vulkan to support 1000s of cross vendor graphics cards (AMD, Qualcomm, NVIDIA & friends). Blazing fast, mobile-enabled, asynchronous and optimized for advanced GPU data processing usecases. Backed by the Linux Foundation.

    C++
    Voir sur GitHub↗2,519
  • kubeflow/mpi-operatorAvatar de kubeflow

    kubeflow/mpi-operator

    528Voir sur GitHub↗

    The MPI Operator makes it easy to run allreduce-style distributed training on Kubernetes. Please check out this blog post for an introduction to MPI Operator and its industry adoption.

    Go
    Voir sur GitHub↗528
  • llnl/rajaAvatar de LLNL

    LLNL/RAJA

    585Voir sur GitHub↗

    comment: # (#################################################################) comment: # (Copyright Lawrence Livermore National Security, LLC and other) comment: # (RAJA Project Developers. See top-level LICENSE and COPYRIGHT) comment: # (files for dates and other details. No copyright…

    C++
    Voir sur GitHub↗585
  • microsoft/deepspeedAvatar de microsoft

    microsoft/DeepSpeed

    42,533Voir sur GitHub↗

    DeepSpeed is a distributed deep learning optimization library and framework designed for the training and inference of massive AI models. It serves as a model parallelism orchestrator and a toolkit for scaling large language models across multiple GPUs and compute nodes. The project distinguishes itself through 3D parallelism orchestration, which combines data, pipeline, and tensor parallelism. It utilizes ZeRO-based memory partitioning to eliminate redundant storage and employs CPU-offload memory management to move weights and optimizer states to system RAM. Additionally, it provides special

    Python
    Voir sur GitHub↗42,533
  • mpi4jax/mpi4jaxAvatar de mpi4jax

    mpi4jax/mpi4jax

    516Voir sur GitHub↗

    mpi4jax

    Pythongpuhigh-performance-computingjax
    Voir sur GitHub↗516
  • nvlabs/cuda-oxideAvatar de NVlabs

    NVlabs/cuda-oxide

    2,825Voir sur GitHub↗

    cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no DSLs, no foreign language bindings, just Rust.

    Rust
    Voir sur GitHub↗2,825
  • openucx/ucxAvatar de openucx

    openucx/ucx

    1,657Voir sur GitHub↗

    Unified Communication X (mailing list - https://elist.ornl.gov/mailman/listinfo/ucx-group)

    Cariescc-plus-plus
    Voir sur GitHub↗1,657
  • raftlib/raftlibAvatar de RaftLib

    RaftLib/RaftLib

    997Voir sur GitHub↗

    RaftLib is an open-source C++ Library that provides a framework for implementing parallel and concurrent data processing pipelines. It is designed to simplify the development of high-performance data processing applications by abstracting away the complexities of parallelism, concurrency, and…

    C++
    Voir sur GitHub↗997
  • rocm-developer-tools/hipAvatar de ROCm-Developer-Tools

    ROCm-Developer-Tools/HIP

    4,362Voir sur GitHub↗

    HIP is a C++ GPU kernel language and cross-platform runtime designed for writing portable high-performance compute applications. It provides a programming interface that allows a single source codebase to execute on both AMD and NVIDIA GPU architectures. The project functions as a compatibility layer that enables the conversion and migration of existing CUDA source code to run on AMD hardware. This is achieved through a syntax mapping that mirrors CUDA and a source-to-source translation process during compilation. The toolkit covers the broader surface of cross-platform GPGPU development, in

    C++
    Voir sur GitHub↗4,362
  • stanfordlegion/legionAvatar de StanfordLegion

    StanfordLegion/legion

    760Voir sur GitHub↗

    Legion is a parallel programming model for distributed, heterogeneous machines.

    C++
    Voir sur GitHub↗760
  • stellar-group/hpxAvatar de STEllAR-GROUP

    STEllAR-GROUP/hpx

    2,858Voir sur GitHub↗

    The C++ Standard Library for Parallelism and Concurrency

    C++
    Voir sur GitHub↗2,858
  • taichi-dev/taichiAvatar de taichi-dev

    taichi-dev/taichi

    27,982Voir sur GitHub↗

    Taichi is a domain-specific programming language embedded in Python designed for high-performance numerical computing and computer graphics. It functions as a parallel compiler that translates high-level mathematical expressions into optimized machine instructions, enabling developers to write compute-intensive algorithms that execute across diverse hardware architectures, including CPUs, GPUs, and specialized accelerators. The project distinguishes itself through a hardware-agnostic execution layer that maps parallel operations to multiple backends such as CUDA, Metal, and Vulkan. By utilizi

    C++computer-graphicsdifferentiable-programminggpu
    Voir sur GitHub↗27,982
  • taskflow/taskflowAvatar de taskflow

    taskflow/taskflow

    12,013Voir sur GitHub↗

    Taskflow is a C++ task-parallel framework designed to build high-performance parallel workflows and complex dependency graphs. It provides a programming model that organizes computational work into directed acyclic graphs, enabling developers to manage concurrency, resource scheduling, and task dependencies across multi-core CPUs and GPU accelerators. The framework distinguishes itself through its ability to orchestrate heterogeneous systems, allowing for the integration of hardware-accelerated kernels and memory operations into unified execution pipelines. It supports dynamic runtime subflow

    C++concurrent-programmingcuda-programminggpu-programming
    Voir sur GitHub↗12,013
  • vosen/zludaAvatar de vosen

    vosen/ZLUDA

    13,945Voir sur GitHub↗

    ZLUDA is a middleware and translation engine designed to enable the execution of unmodified proprietary compute binaries on non-native graphics hardware. It functions as a compatibility layer that bridges vendor-specific compute interfaces with open standards, allowing software originally restricted to a single hardware ecosystem to operate on alternative graphics processing units. The project achieves this through a combination of dynamic library interception and runtime instruction translation. By replacing standard system libraries and mapping proprietary compute calls to open standards, t

    Rustcudarust
    Voir sur GitHub↗13,945