awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
inferflow avatar

inferflow/inferflow

0
View on GitHub↗
251 星标·24 分支·C++·MIT·6 次浏览

Inferflow

Inferflow is an efficient and highly configurable inference engine for large language models (LLMs).

Features

  • Inference Frameworks - Highly configurable engine for efficient model execution.

Star 历史

inferflow/inferflow 的 Star 历史图表inferflow/inferflow 的 Star 历史图表

AI 搜索

探索更多 awesome 仓库

用简单的语言描述您的需求 —— AI 将根据相关性为您从数千个精选开源项目中进行排序。

Start searching with AI

Inferflow 的开源替代方案

相似的开源项目,按与 Inferflow 的功能重合度排序。
  • efeslab/nanoflowefeslab 的头像

    efeslab/Nanoflow

    965在 GitHub 上查看↗

    A throughput-oriented high-performance serving framework for LLMs

    Jupyter Notebookcudainferencellama2
    在 GitHub 上查看↗965
  • flashinfer-ai/flashinferflashinfer-ai 的头像

    flashinfer-ai/flashinfer

    4,996在 GitHub 上查看↗

    FlashInfer is a library of high-performance GPU kernels purpose-built for accelerating large language model inference. It provides optimized implementations for attention operations (including flash attention, page attention, multi-head latent attention, and cascade attention) using paged key-value caches, fused kernel composition, and just-in-time compilation. The library also includes specialized kernels for mixture-of-experts layers, block-scaled low-precision quantization (FP8, FP4), and distributed collective communication. What distinguishes FlashInfer is its fused all-reduce communicat

    Pythonattentioncudadistributed-inference
    在 GitHub 上查看↗4,996
  • flexflow/flexflowflexflow 的头像

    flexflow/FlexFlow

    1,889在 GitHub 上查看↗

    Automatically Discovering Fast Parallelization Strategies for Distributed Deep Neural Network Training

    C++
    在 GitHub 上查看↗1,889
  • bigscience-workshop/petalsbigscience-workshop 的头像

    bigscience-workshop/petals

    10,208在 GitHub 上查看↗

    Petals is a decentralized framework and inference engine for running large language models across a peer-to-peer network. It enables the execution of models that exceed the memory of any single machine by splitting computations and model layers across a collaborative swarm of GPUs. The system functions as a collaborative compute network where participants share local GPU resources and host model weights. It supports distributed prompt-tuning to adapt massive models to specific tasks and allows for the establishment of private compute swarms to process sensitive data within restricted, trusted

    Python
    在 GitHub 上查看↗10,208
查看 Inferflow 的所有 19 个替代方案→

常见问题解答

inferflow/inferflow 是做什么的?

Inferflow is an efficient and highly configurable inference engine for large language models (LLMs).

inferflow/inferflow 的主要功能有哪些?

inferflow/inferflow 的主要功能包括:Inference Frameworks。

inferflow/inferflow 有哪些开源替代品?

inferflow/inferflow 的开源替代品包括: efeslab/nanoflow — A throughput-oriented high-performance serving framework for LLMs. flashinfer-ai/flashinfer — FlashInfer is a library of high-performance GPU kernels purpose-built for accelerating large language model inference.… flexflow/flexflow — Automatically Discovering Fast Parallelization Strategies for Distributed Deep Neural Network Training. fminference/flexgen — FlexGen is an inference engine for large language models designed for high-throughput execution on single or multiple… ggerganov/llama.cpp — llama.cpp is a high-performance C++ inference engine and runtime for executing large language models locally across… bigscience-workshop/petals — Petals is a decentralized framework and inference engine for running large language models across a peer-to-peer…