awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
flexflow avatar

flexflow/FlexFlow

0
View on GitHub↗
1,889 stars·252 forks·C++·Apache-2.0·7 viewsflexflow.ai↗

FlexFlow

Automatically Discovering Fast Parallelization Strategies for Distributed Deep Neural Network Training

Features

  • Inference Frameworks - Framework for speculative inference and token verification.

Star history

Star history chart for flexflow/flexflowStar history chart for flexflow/flexflow

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to FlexFlow

Similar open-source projects, ranked by how many features they share with FlexFlow.
  • efeslab/nanoflowefeslab avatar

    efeslab/Nanoflow

    965View on GitHub↗

    A throughput-oriented high-performance serving framework for LLMs

    Jupyter Notebookcudainferencellama2
    View on GitHub↗965
  • flashinfer-ai/flashinferflashinfer-ai avatar

    flashinfer-ai/flashinfer

    4,996View on GitHub↗

    FlashInfer is a library of high-performance GPU kernels purpose-built for accelerating large language model inference. It provides optimized implementations for attention operations (including flash attention, page attention, multi-head latent attention, and cascade attention) using paged key-value caches, fused kernel composition, and just-in-time compilation. The library also includes specialized kernels for mixture-of-experts layers, block-scaled low-precision quantization (FP8, FP4), and distributed collective communication. What distinguishes FlashInfer is its fused all-reduce communicat

    Pythonattentioncudadistributed-inference
    View on GitHub↗4,996
  • fminference/flexgenFMInference avatar

    FMInference/FlexGen

    9,366View on GitHub↗

    FlexGen is an inference engine for large language models designed for high-throughput execution on single or multiple GPUs. It functions as a framework for managing model execution through a combination of memory offloading, weight compression, and pipeline orchestration. The system enables the execution of models that exceed available GPU memory by moving tensors and caches between GPU memory, system RAM, and disk storage. It utilizes 4-bit weight quantization to reduce the memory footprint of model parameters, allowing for increased batch processing capacity. The project covers distributed

    Python
    View on GitHub↗9,366
  • bigscience-workshop/petalsbigscience-workshop avatar

    bigscience-workshop/petals

    10,208View on GitHub↗

    Petals is a decentralized framework and inference engine for running large language models across a peer-to-peer network. It enables the execution of models that exceed the memory of any single machine by splitting computations and model layers across a collaborative swarm of GPUs. The system functions as a collaborative compute network where participants share local GPU resources and host model weights. It supports distributed prompt-tuning to adapt massive models to specific tasks and allows for the establishment of private compute swarms to process sensitive data within restricted, trusted

    Python
    View on GitHub↗10,208
See all 19 alternatives to FlexFlow→

Frequently asked questions

What does flexflow/flexflow do?

Automatically Discovering Fast Parallelization Strategies for Distributed Deep Neural Network Training

What are the main features of flexflow/flexflow?

The main features of flexflow/flexflow are: Inference Frameworks.

What are some open-source alternatives to flexflow/flexflow?

Open-source alternatives to flexflow/flexflow include: efeslab/nanoflow — A throughput-oriented high-performance serving framework for LLMs. flashinfer-ai/flashinfer — FlashInfer is a library of high-performance GPU kernels purpose-built for accelerating large language model inference.… fminference/flexgen — FlexGen is an inference engine for large language models designed for high-throughput execution on single or multiple… ggerganov/llama.cpp — llama.cpp is a high-performance C++ inference engine and runtime for executing large language models locally across… inferflow/inferflow — Inferflow is an efficient and highly configurable inference engine for large language models (LLMs). bigscience-workshop/petals — Petals is a decentralized framework and inference engine for running large language models across a peer-to-peer…