awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to nvidia/libcudacxx

Projects sharing features with Libcudacxx

30 open-source projects similar to nvidia/libcudacxx, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • ddemidov/vexclddemidov avatar

    ddemidov/vexcl

    721View on GitHub↗

    VexCL is a C++ vector expression template library for OpenCL/CUDA/OpenMP

    C++
    View on GitHub↗721
  • arrayfire/arrayfirearrayfire avatar

    arrayfire/arrayfire

    4,888View on GitHub↗

    ArrayFire is a hardware-agnostic compute framework and JIT-compiled tensor engine designed for high-performance numerical computing. It serves as a GPU numerical computing library and parallel signal processing toolkit that abstracts hardware backends, allowing the same codebase to execute across various GPU architectures and CPUs. The project distinguishes itself through a JIT engine that uses expression compilation to fuse operations and minimize memory overhead. It employs a deferred execution graph to optimize computation chains and provides interoperability primitives to share data and e

    C++arrayfirecc-plus-plus
    View on GitHub↗4,888
  • nvidia/tensorrtNVIDIA avatar

    NVIDIA/TensorRT

    13,076View on GitHub↗

    TensorRT is a deep learning inference engine and software development kit designed to optimize and deploy neural networks for high-performance execution on NVIDIA GPUs. It functions as a GPU acceleration framework that reduces latency and increases throughput for trained models during production deployment. The toolkit imports models from the Open Neural Network Exchange format and transforms them into optimized engines. It utilizes graph-based model optimization, layer-fusion kernel generation, and precision-based quantization to convert floating point weights into lower precision formats.

    C++deep-learninggpu-accelerationinference
    View on GitHub↗13,076
  • xilinx/pynqXilinx avatar

    Xilinx/PYNQ

    2,311View on GitHub↗

    Python Productivity for ZYNQ

    Jupyter Notebook
    View on GitHub↗2,311

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • dask/daskdask avatar

    dask/dask

    13,746View on GitHub↗

    Dask is a parallel computing framework and distributed task scheduler designed to scale Python data science workflows from single machines to large clusters. It functions as a cluster resource manager that orchestrates computational logic by representing tasks and their dependencies as directed acyclic graphs. This architecture allows the system to automate the distribution of workloads across available hardware while managing complex execution requirements. The project distinguishes itself through a lazy evaluation engine that defers data operations until they are explicitly requested, enabl

    Pythondasknumpypandas
    View on GitHub↗13,746
  • cupy/cupycupy avatar

    cupy/cupy

    11,000View on GitHub↗

    CuPy is a CUDA array computing library that implements a NumPy-compatible interface for executing array operations and numerical computing on NVIDIA GPUs. It serves as a GPU-accelerated numerical library and a CUDA-based SciPy implementation, offloading heavy calculations to graphics hardware to increase processing speed for scientific and engineering workloads. The library enables multi-framework tensor exchange, allowing data buffers to be shared between different deep learning frameworks using standardized memory layouts to avoid memory copies. It also supports custom GPU kernel integratio

    Python
    View on GitHub↗11,000
  • numba/numbanumba avatar

    numba/numba

    10,918View on GitHub↗

    Numba is a just-in-time compiler that translates high-level Python functions into optimized machine code at runtime. By leveraging the LLVM compiler infrastructure, it provides a framework for accelerating numerical data processing and mathematical computations, enabling performance levels comparable to statically compiled languages. The project distinguishes itself through its ability to perform type-inference-based specialization, which generates machine instructions tailored to the specific data types used during execution. It employs a lazy compilation pipeline that defers translation unt

    Pythoncompilercudallvm
    View on GitHub↗10,918
  • thrust/thrustthrust avatar

    thrust/thrust

    5,003View on GitHub↗

    Thrust is a heterogeneous computing library and C++ template library that provides a collection of high-level templates for executing data-parallel operations. It functions as a parallel algorithms library designed to work across different hardware backends, including multicore CPUs and NVIDIA GPU hardware. The framework utilizes a header-only implementation and a generic-programming policy interface to abstract the differences between CPU and GPU memory and execution models. It employs an iterator-based data abstraction to provide a uniform interface for accessing elements across host RAM an

    C++
    View on GitHub↗5,003
  • microsoft/rusttrainingmicrosoft avatar

    microsoft/RustTraining

    14,636View on GitHub↗

    This project is a structured Rust programming curriculum and systems programming course designed to take learners from beginner to expert levels. It provides a comprehensive set of training materials focused on mastering the core syntax, idioms, and technical foundations of the Rust language. The project features a specialized language transition framework that maps concepts from C++, managed languages, and dynamic typing to Rust idioms. This allows developers from different ecosystems to translate architectural patterns and memory models into idiomatic Rust. The training covers a broad rang

    Rust
    View on GitHub↗14,636
  • atgreen/cl-cancelatgreen avatar

    atgreen/cl-cancel

    5View on GitHub↗

    Cancellation propagation library for Common Lisp with deadlines and timeouts

    Common Lisp
    View on GitHub↗5
  • boostorg/computeboostorg avatar

    boostorg/compute

    1,654View on GitHub↗

    A C++ GPU Computing Library for OpenCL

    C++boostc-plus-pluscompute
    View on GitHub↗1,654
  • bloomen/transwarpbloomen avatar

    bloomen/transwarp

    633View on GitHub↗

    A header-only C++ library for task concurrency

    C++
    View on GitHub↗633
  • borodust/cl-flowborodust avatar

    borodust/cl-flow

    52View on GitHub↗

    Reactive computation tree library for non-blocking concurrent Common Lisp

    Common Lisp
    View on GitHub↗52
  • brown/swank-crewbrown avatar

    brown/swank-crew

    48View on GitHub↗

    Common Lisp distributed computation framework implemented using Swank Client

    Common Lisp
    View on GitHub↗48
  • amanieu/asyncplusplusAmanieu avatar

    Amanieu/asyncplusplus

    1,417View on GitHub↗

    Async++ concurrency framework for C++11

    C++
    View on GitHub↗1,417
  • bloomberg/quantumbloomberg avatar

    bloomberg/quantum

    631View on GitHub↗

    Powerful multi-threaded coroutine dispatcher and parallel execution engine

    C++
    View on GitHub↗631
  • computationalradiationphysics/cuplaComputationalRadiationPhysics avatar

    ComputationalRadiationPhysics/cupla

    4View on GitHub↗

    The project alpaka has moved to https://github.com/alpaka-group/cupla

    View on GitHub↗4
  • basiliscos/cpp-rotorbasiliscos avatar

    basiliscos/cpp-rotor

    388View on GitHub↗

    Event loop friendly C++ actor micro-framework, supervisable

    C++
    View on GitHub↗388
  • apple/swift-corelibs-libdispatchapple avatar

    apple/swift-corelibs-libdispatch

    2,596View on GitHub↗

    The libdispatch Project, (a.k.a. Grand Central Dispatch), for concurrency on multicore hardware

    C
    View on GitHub↗2,596
  • concurrencykit/ckconcurrencykit avatar

    concurrencykit/ck

    2,650View on GitHub↗

    Concurrency primitives, safe memory reclamation mechanisms and non-blocking (including lock-free) data structures designed to aid in the research, design and implementation of high performance concurrent systems developed in C99+.

    C
    View on GitHub↗2,650
  • conorwilliams/libforkConorWilliams avatar

    ConorWilliams/libfork

    879View on GitHub↗

    A bleeding-edge, lock-free, wait-free, continuation-stealing tasking library built on C++20's coroutines

    C++
    View on GitHub↗879
  • cosmos72/stmxcosmos72 avatar

    cosmos72/stmx

    258View on GitHub↗

    High performance Transactional Memory for Common Lisp

    Common Lisp
    View on GitHub↗258
  • computationalradiationphysics/alpakaComputationalRadiationPhysics avatar

    ComputationalRadiationPhysics/alpaka

    4View on GitHub↗

    The project alpaka has moved to https://github.com/alpaka-group/alpaka

    View on GitHub↗4
  • cameron314/readerwriterqueuecameron314 avatar

    cameron314/readerwriterqueue

    4,576View on GitHub↗

    This project is a single-producer single-consumer concurrent queue for C++ designed for lock-free data exchange between threads. It provides a thread-safe mechanism to transfer data without the use of mutexes or locks. The queue is implemented as a contiguous circular buffer that supports dynamic capacity growth to prevent data loss when the queue reaches its limit. It utilizes atomic synchronization and wait-free index management to coordinate data access between the writing and reading threads. The library covers inter-thread communication and buffer management, offering both blocking and

    C++
    View on GitHub↗4,576
  • david-haim/concurrencppDavid-Haim avatar

    David-Haim/concurrencpp

    2,755View on GitHub↗

    Modern concurrency for C++. Tasks, executors, timers and C++20 coroutines to rule them all

    C++
    View on GitHub↗2,755
  • atgreen/cl-natsatgreen avatar

    atgreen/cl-nats

    11View on GitHub↗

    A full-featured NATS messaging client for Common Lisp.

    Common Lisp
    View on GitHub↗11
  • digital-fabric/polyphonydigital-fabric avatar

    digital-fabric/polyphony

    662View on GitHub↗

    Fine-grained concurrency for Ruby

    C
    View on GitHub↗662
  • eventmachine/eventmachineeventmachine avatar

    eventmachine/eventmachine

    4,283View on GitHub↗

    EventMachine is a reactor-pattern network framework for Ruby that provides an asynchronous I/O library for performing non-blocking network and file operations. It functions as a network server framework used to build scalable TCP and UDP servers and clients that process multiple simultaneous requests. The framework implements a concurrency model that dispatches network events to registered handlers using a single-threaded event loop. This approach allows for the management of high-concurrency network connections without the overhead of multi-threaded programming. The library covers the devel

    Ruby
    View on GitHub↗4,283
  • eyalroz/cuda-api-wrapperseyalroz avatar

    eyalroz/cuda-api-wrappers

    890View on GitHub↗

    Thin, unified, C++-flavored wrappers for the CUDA APIs

    C++
    View on GitHub↗890
  • cameron314/concurrentqueuecameron314 avatar

    cameron314/concurrentqueue

    12,070View on GitHub↗

    ConcurrentQueue is a header-only C++ template library that provides a lock-free data structure for multi-producer multi-consumer thread communication. It functions as a synchronization primitive designed to coordinate data flow between concurrent execution units using atomic operations rather than traditional mutex locking. The library distinguishes itself through a design that minimizes contention and synchronization overhead. It utilizes sub-queue token mapping to distribute workloads across partitioned internal queues and supports bulk operations to transfer multiple data elements in singl

    C++
    View on GitHub↗12,070