8 dépôts
Tools for analyzing execution timing to optimize the speed and efficiency of model inference.
Distinct from Model Performance Benchmarks: Focuses on real-time inference speed tuning rather than comparative training and ranking of multiple candidate models
Explore 8 awesome GitHub repositories matching artificial intelligence & ml · Inference Speed Profiling. Refine with filters or upvote what's useful.
jetson-inference is a set of libraries and tools for executing optimized deep learning models on embedded GPU hardware. Its primary purpose is to enable real-time computer vision and AI inference at the edge with low latency and high throughput. The project distinguishes itself through high-performance streaming analytics and the ability to execute concurrent AI pipelines on auto-grade silicon. It provides specialized support for multi-sensor stream processing, utilizing zero-copy data transport to load camera frames directly into GPU memory. The codebase covers a broad surface of capabiliti
Profiles model performance and analyzes execution timing to tune inference speed and efficiency.
MMPose is a PyTorch-based pose estimation toolbox and deep learning training pipeline designed for detecting 2D and 3D keypoints on humans, animals, and faces. It serves as a computer vision model zoo and a framework for both 2D pose estimation and 3D pose lifting. The project is distinguished by its modular architecture and extensibility, employing a registry-based system and hierarchical configurations to allow for custom algorithm integration and model pipeline customization. It supports diverse estimation paradigms, including top-down, bottom-up, and two-stage pose lifting workflows. The
Measures execution speed and frames per second of deployed models using representative test images.
GLM-4 is a large language model and fine-tuning framework designed for human-like text production, complex reasoning, and multilingual conversation. It functions as a multimodal system capable of processing high-resolution visual content and as a long-context model designed to analyze documents with a context window of up to one million tokens. The project differentiates itself through a function calling interface that enables AI agent development by connecting the model to external APIs and real-time web browsing. It includes specialized capabilities for generating functional programming cod
Calculates tokens-per-second performance on local hardware to profile and track inference speed efficiency.
LMCache is a distributed key-value cache manager and tiering system designed to accelerate large language model inference. It functions as a tiered storage layer that offloads tensors from GPU memory to CPU RAM, local disks, or remote object stores, enabling the reuse of cached prefixes across different inference sessions and serving engines. The system differentiates itself through a disaggregated prefill-decode model, which separates prompt processing from token generation by transferring caches between distributed compute nodes. It utilizes peer-to-peer orchestration to share and retrieve
Simulates configurable traffic patterns to report speed and throughput metrics for the inference engine.
Dynamo is a distributed inference orchestration platform designed for large language models. It functions as a system to coordinate prefill and decode phases across GPU nodes, utilizing a multi-backend runtime adapter to connect engines like vLLM and TensorRT-LLM through a unified block-oriented memory interface. An OpenAI-compatible API server provides the frontend for integration with existing tools and clients. The project is distinguished by its disaggregated serving architecture, which separates prompt processing and token generation onto independent GPU pools to optimize throughput and
Mimics backend API behavior and synthetic traffic patterns to validate routing and infrastructure logic without consuming GPUs.
Ce projet est un optimiseur de performance et un benchmark de ressources pour AWS Lambda. Il analyse le compromis entre la vitesse d'exécution et le coût en testant diverses configurations de mémoire pour identifier les paramètres les plus rentables et minimiser les dépenses opérationnelles. L'outil utilise un orchestrateur AWS Step Functions pour automatiser l'exécution et la collecte de données de plusieurs tests de fonctions à différents niveaux de puissance. Il simule des charges de travail de production en injectant des données statiques ou distantes personnalisées et en utilisant une distribution de charge utile pondérée pour imiter des modèles de trafic réels. La suite couvre plusieurs domaines de capacité, incluant l'échantillonnage itératif de la mémoire et la modélisation des coûts basée sur les métriques pour visualiser les compromis de performance. Elle fournit un nettoyage automatisé des ressources pour les versions et alias de fonctions temporaires, une configuration réseau privée pour les ressources internes restreintes et le chargement de charge utile à distance pour contourner les limites de taille d'invocation standard. Le déploiement est géré via des constructions d'infrastructure-as-code pour garantir une configuration d'environnement cohérente et une répétabilité.
Simulates production traffic by distributing test input payloads based on assigned relative probability weights.
LightLLM is a high-performance serving framework for deploying and executing large language models. It functions as a multi-GPU inference engine and server capable of handling dense architectures, mixture-of-experts designs, and multimodal models that process both text and images. The system is distinguished by its specialized support for Mixture-of-Experts models using expert parallelism and fused kernels. It implements structured text generation through deterministic state machines and pushdown automata to enforce precise output formats. To optimize throughput, the framework employs specula
Includes detailed profiling for prefill and decode stage throughput and latency across multi-GPU configurations.
This project is a containerized local AI infrastructure stack designed to deploy large language models and vector databases on private hardware. It functions as an orchestration platform that combines AI runners, knowledge graphs, and a visual workflow builder for creating agentic chatflows and automating tasks via tool integration. The platform distinguishes itself through a low-code approach to agent orchestration, utilizing a visual interface to design complex sequences and connect agents to external tools and search engines. It includes a dedicated local observability stack to track promp
Optimizes model processing speed by selecting hardware-specific configuration profiles for GPUs or CPUs.