awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

8 个仓库

Awesome GitHub RepositoriesInference Speed Profiling

Tools for analyzing execution timing to optimize the speed and efficiency of model inference.

Distinct from Model Performance Benchmarks: Focuses on real-time inference speed tuning rather than comparative training and ranking of multiple candidate models

Explore 8 awesome GitHub repositories matching artificial intelligence & ml · Inference Speed Profiling. Refine with filters or upvote what's useful.

Awesome Inference Speed Profiling GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • dusty-nv/jetson-inferencedusty-nv 的头像

    dusty-nv/jetson-inference

    8,734在 GitHub 上查看↗

    jetson-inference is a set of libraries and tools for executing optimized deep learning models on embedded GPU hardware. Its primary purpose is to enable real-time computer vision and AI inference at the edge with low latency and high throughput. The project distinguishes itself through high-performance streaming analytics and the ability to execute concurrent AI pipelines on auto-grade silicon. It provides specialized support for multi-sensor stream processing, utilizing zero-copy data transport to load camera frames directly into GPU memory. The codebase covers a broad surface of capabiliti

    Profiles model performance and analyzes execution timing to tune inference speed and efficiency.

    C++caffecomputer-visiondeep-learning
    在 GitHub 上查看↗8,734
  • open-mmlab/mmposeopen-mmlab 的头像

    open-mmlab/mmpose

    7,374在 GitHub 上查看↗

    MMPose is a PyTorch-based pose estimation toolbox and deep learning training pipeline designed for detecting 2D and 3D keypoints on humans, animals, and faces. It serves as a computer vision model zoo and a framework for both 2D pose estimation and 3D pose lifting. The project is distinguished by its modular architecture and extensibility, employing a registry-based system and hierarchical configurations to allow for custom algorithm integration and model pipeline customization. It supports diverse estimation paradigms, including top-down, bottom-up, and two-stage pose lifting workflows. The

    Measures execution speed and frames per second of deployed models using representative test images.

    Pythonanimal-pose-estimationbenchmarkcpm
    在 GitHub 上查看↗7,374
  • zai-org/glm-4zai-org 的头像

    zai-org/GLM-4

    7,058在 GitHub 上查看↗

    GLM-4 is a large language model and fine-tuning framework designed for human-like text production, complex reasoning, and multilingual conversation. It functions as a multimodal system capable of processing high-resolution visual content and as a long-context model designed to analyze documents with a context window of up to one million tokens. The project differentiates itself through a function calling interface that enables AI agent development by connecting the model to external APIs and real-time web browsing. It includes specialized capabilities for generating functional programming cod

    Calculates tokens-per-second performance on local hardware to profile and track inference speed efficiency.

    Pythonchatglmchatglm-6bglm
    在 GitHub 上查看↗7,058
  • lmcache/lmcacheLMCache 的头像

    LMCache/LMCache

    6,909在 GitHub 上查看↗

    LMCache is a distributed key-value cache manager and tiering system designed to accelerate large language model inference. It functions as a tiered storage layer that offloads tensors from GPU memory to CPU RAM, local disks, or remote object stores, enabling the reuse of cached prefixes across different inference sessions and serving engines. The system differentiates itself through a disaggregated prefill-decode model, which separates prompt processing from token generation by transferring caches between distributed compute nodes. It utilizes peer-to-peer orchestration to share and retrieve

    Simulates configurable traffic patterns to report speed and throughput metrics for the inference engine.

    Pythonamdcudafast
    在 GitHub 上查看↗6,909
  • ai-dynamo/dynamoai-dynamo 的头像

    ai-dynamo/dynamo

    6,112在 GitHub 上查看↗

    Dynamo is a distributed inference orchestration platform designed for large language models. It functions as a system to coordinate prefill and decode phases across GPU nodes, utilizing a multi-backend runtime adapter to connect engines like vLLM and TensorRT-LLM through a unified block-oriented memory interface. An OpenAI-compatible API server provides the frontend for integration with existing tools and clients. The project is distinguished by its disaggregated serving architecture, which separates prompt processing and token generation onto independent GPU pools to optimize throughput and

    Mimics backend API behavior and synthetic traffic patterns to validate routing and infrastructure logic without consuming GPUs.

    Rust
    在 GitHub 上查看↗6,112
  • alexcasalboni/aws-lambda-power-tuningalexcasalboni 的头像

    alexcasalboni/aws-lambda-power-tuning

    6,028在 GitHub 上查看↗

    该项目是 AWS Lambda 的性能优化器和资源基准测试工具。它通过测试各种内存配置来分析执行速度与成本之间的权衡,以确定最具成本效益的设置并最小化运营支出。 该工具利用 AWS Step Functions 编排器来自动化执行和收集跨不同功率级别的多个函数测试运行的数据。它通过注入自定义静态或远程数据并使用加权有效负载分布来模拟生产工作负载,以模仿真实世界的流量模式。 该套件涵盖了几个功能领域,包括迭代内存采样和基于指标的成本建模,以可视化性能权衡。它为临时函数版本和别名提供自动化资源清理、为受限内部资源提供私有网络配置,以及提供远程有效负载加载以绕过标准调用大小限制。 部署通过基础设施即代码(IaC)结构进行处理,以确保环境设置的一致性和可重复性。

    Simulates production traffic by distributing test input payloads based on assigned relative probability weights.

    JavaScript
    在 GitHub 上查看↗6,028
  • modeltc/lightllmModelTC 的头像

    ModelTC/LightLLM

    3,901在 GitHub 上查看↗

    LightLLM is a high-performance serving framework for deploying and executing large language models. It functions as a multi-GPU inference engine and server capable of handling dense architectures, mixture-of-experts designs, and multimodal models that process both text and images. The system is distinguished by its specialized support for Mixture-of-Experts models using expert parallelism and fused kernels. It implements structured text generation through deterministic state machines and pushdown automata to enforce precise output formats. To optimize throughput, the framework employs specula

    Includes detailed profiling for prefill and decode stage throughput and latency across multi-GPU configurations.

    Pythondeep-learninggptllama
    在 GitHub 上查看↗3,901
  • coleam00/local-ai-packagedcoleam00 的头像

    coleam00/local-ai-packaged

    3,539在 GitHub 上查看↗

    This project is a containerized local AI infrastructure stack designed to deploy large language models and vector databases on private hardware. It functions as an orchestration platform that combines AI runners, knowledge graphs, and a visual workflow builder for creating agentic chatflows and automating tasks via tool integration. The platform distinguishes itself through a low-code approach to agent orchestration, utilizing a visual interface to design complex sequences and connect agents to external tools and search engines. It includes a dedicated local observability stack to track promp

    Optimizes model processing speed by selecting hardware-specific configuration profiles for GPUs or CPUs.

    Python
    在 GitHub 上查看↗3,539
  1. Home
  2. Artificial Intelligence & ML
  3. Cross-Model Comparators
  4. Model Performance Benchmarks
  5. Inference Speed Profiling

探索子标签

  • Hardware Configuration ProfilesPredefined performance settings tailored to specific CPU or GPU hardware to optimize inference speed. **Distinct from Inference Speed Profiling:** Focuses on using profiles to set optimization levels rather than just analyzing execution timing.
  • Workload Simulations1 个子标签Tools for simulating specific traffic patterns to measure inference throughput and speed. **Distinct from Inference Speed Profiling:** Focuses on synthetic traffic generation and load testing rather than just timing profiling.