3 Repos
Tools for simulating specific traffic patterns to measure inference throughput and speed.
Distinct from Inference Speed Profiling: Focuses on synthetic traffic generation and load testing rather than just timing profiling.
Explore 3 awesome GitHub repositories matching artificial intelligence & ml · Workload Simulations. Refine with filters or upvote what's useful.
LMCache is a distributed key-value cache manager and tiering system designed to accelerate large language model inference. It functions as a tiered storage layer that offloads tensors from GPU memory to CPU RAM, local disks, or remote object stores, enabling the reuse of cached prefixes across different inference sessions and serving engines. The system differentiates itself through a disaggregated prefill-decode model, which separates prompt processing from token generation by transferring caches between distributed compute nodes. It utilizes peer-to-peer orchestration to share and retrieve
Simulates configurable traffic patterns to report speed and throughput metrics for the inference engine.
Dynamo is a distributed inference orchestration platform designed for large language models. It functions as a system to coordinate prefill and decode phases across GPU nodes, utilizing a multi-backend runtime adapter to connect engines like vLLM and TensorRT-LLM through a unified block-oriented memory interface. An OpenAI-compatible API server provides the frontend for integration with existing tools and clients. The project is distinguished by its disaggregated serving architecture, which separates prompt processing and token generation onto independent GPU pools to optimize throughput and
Mimics backend API behavior and synthetic traffic patterns to validate routing and infrastructure logic without consuming GPUs.
Dieses Projekt ist ein Leistungsoptimierer und Ressourcen-Bencher für AWS Lambda. Es analysiert das Verhältnis zwischen Ausführungsgeschwindigkeit und Kosten, indem es verschiedene Speicherkonfigurationen testet, um die kosteneffizientesten Einstellungen zu identifizieren und die Betriebsausgaben zu minimieren. Das Tool nutzt einen AWS-Step-Functions-Orchestrator, um die Ausführung und Datensammlung mehrerer Funktionstestläufe über verschiedene Leistungsstufen hinweg zu automatisieren. Es simuliert Produktions-Workloads durch das Injizieren benutzerdefinierter statischer oder Remote-Daten und die Verwendung gewichteter Payload-Verteilung, um reale Verkehrsmuster nachzuahmen. Die Suite deckt mehrere Funktionsbereiche ab, einschließlich iterativer Speicherabtastung und metrikbasierter Kostenmodellierung zur Visualisierung von Leistungs-Trade-offs. Sie bietet automatisierte Ressourcenbereinigung für temporäre Funktionsversionen und Aliase, private Netzwerkkonfiguration für eingeschränkte interne Ressourcen und Remote-Payload-Laden, um Standard-Aufrufgrößenbeschränkungen zu umgehen. Die Bereitstellung erfolgt über Infrastructure-as-Code-Konstrukte, um eine konsistente Umgebungseinrichtung und Wiederholbarkeit zu gewährleisten.
Simulates production traffic by distributing test input payloads based on assigned relative probability weights.