awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
kserve avatar

kserve/kserve

0
View on GitHub↗

Kserve

KServe is a Kubernetes-native platform for deploying and serving machine learning models as scalable inference services. It supports both generative AI models, including large language models, and traditional predictive models from frameworks such as TensorFlow, PyTorch, Scikit-Learn, XGBoost, and ONNX. The platform manages the full lifecycle of model deployments, including revision tracking, canary rollouts, A/B testing, and automatic rollbacks, and provides serverless scale-to-zero capabilities for cost-efficient resource management.

KServe distinguishes itself through a standardized inference protocol that supports REST, gRPC, and streaming responses, along with an OpenAI-compatible API for seamless integration with existing LLM tooling. It offers advanced traffic management with canary deployments, weighted traffic splitting, and inference graph routing for orchestrating complex multi-model pipelines and ensembles. The platform also provides GPU-accelerated inference, disaggregated prefill-decode serving, and local model caching to optimize performance and reduce latency for large models.

Beyond core serving, KServe includes capabilities for model monitoring and observability, detecting input drift and outliers, and collecting inference metrics with Prometheus, OpenTelemetry, and payload logging. It supports authentication and access control for inference endpoints, including API key, OAuth2, and AWS request signing. The platform integrates with Kubeflow Pipelines for end-to-end ML workflows and offers flexible installation via Helm charts or Kustomize overlays.

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Start searching with AI
kserve.github.io/website
↗

Features

  • Kubernetes ML Platforms - Deploys and manages machine learning models as scalable inference services on Kubernetes with automatic scaling and traffic routing.
  • OpenAI-Compatible APIs - Accepts inference requests using the OpenAI protocol for seamless integration with existing LLM tooling.
  • KServe - Routes LLM inference requests through a dedicated AI gateway for high-performance networking and workload management.
  • Model Request Routing - Directs API requests to the appropriate generative AI service based on model identifiers in the request payload.
  • Multi-Model Pipeline Routings - Coordinates complex inference workflows with multi-model pipelines, conditional logic, and request chaining.
  • AI Request Routing - Manages rate-limiting based on tokens and routes requests to different models through a unified API for generative tasks.
  • Generative AI Model Serving - Serves large language models with OpenAI-compatible APIs, streaming, and optimized memory management.
  • GPU-Accelerated Inference - Optimizes inference with GPU acceleration, KV cache offloading, and local model caching for low-latency LLM serving.
  • Standardized Inference Protocols - Standardizes inference requests and responses across REST and gRPC with health checking and metadata endpoints.
  • Standardized V2 Protocols - Performs inference using a standardized V2 protocol that supports REST and gRPC, health checking, and rich metadata.
  • Inference Scaling - Adjusts the number of running model replicas based on incoming request traffic to maintain performance and reduce idle cost.
  • Automatic Inference Workload Scaling - Automatically scales inference serving replicas based on token throughput and GPU utilization metrics.
  • Large Language Model Serving - Deploys large language models as scalable inference services on Kubernetes.
  • High-Throughput Text Inference - Accelerates large language model inference using a specialized runtime optimized for throughput and latency.
  • Multi-Node Distributions - Distributes model serving across multiple nodes and GPUs to achieve high throughput and low latency for large models.
  • Registry Artifact Fetchers - Fetches model artifacts from Hugging Face Hub or MLflow registries for deployment.
  • Model Lifecycle Management - Manages model deployments with revision tracking, canary rollouts, A/B testing, and automatic rollbacks on Kubernetes.
  • Model Inference Execution - Sends prediction requests to deployed models via HTTP, DNS, or Python client and receives model outputs.
  • Model Serving & Deployment - Serves traditional ML models from TensorFlow, PyTorch, Scikit-Learn, and ONNX with real-time and batch inference.
  • Hugging Face Deployments - Deploys Hugging Face Transformers models as scalable inference endpoints on Kubernetes.
  • Kubernetes Custom Resource Deployments - Defines model serving workloads as Kubernetes custom resources for declarative lifecycle management.
  • Multi-Framework Model Serving - Supports serving models from TensorFlow, PyTorch, Scikit-Learn, XGBoost, ONNX, and Hugging Face with standardized inference protocols.
  • Model Serving Runtimes - Loads model-specific serving containers dynamically based on model format and runtime declarations.
  • Multi-Model Workflow Coordinators - Orchestrates ensembles and pipelines of models using inference graphs for complex multi-step predictions.
  • Large Language Model Deployments - Deploys large language models as scalable inference services using an optimized vLLM backend.
  • Automatic Serving Scaling - Automatically scales model serving instances based on request volume, including scaling to zero when idle.
  • Model Artifact Loaders - Loads model artifacts from S3, GCS, or Azure Blob storage during deployment.
  • Unified Model Loaders - Loads models from cloud storage, Hugging Face, or local volumes through a unified abstraction.
  • Serving Workload Scaling - Scales model serving replicas up or down based on request volume, including scale-to-zero for cost efficiency.
  • Knative-Based Deployments - Leverages Knative for automatic scaling of inference services, including scale-to-zero.
  • GPU-Accelerated Deployments - Deploys models with full resource control and GPU support for both generative and predictive inference workloads.
  • Model Rollout Orchestrators - Coordinates model version deployments, monitors health, and reverts to stable versions when needed.
  • Custom Resource Controllers - Manages model lifecycle through Kubernetes custom resources that reconcile desired state with cluster state.
  • Canary Deployment Controllers - Rolls out new model versions gradually using canary deployments and A/B testing to reduce risk.
  • Direct Deployments - Deploys models directly from the Hugging Face Hub with native support and streamlined configuration.
  • Serving Runtimes - Provides a specialized runtime for deploying Hugging Face transformer models as scalable inference services.
  • Model Deployment Management - Manages the full lifecycle of model deployments with revision tracking, canary rollouts, and A/B testing.
  • Model Serving - Deploys models from TensorFlow, PyTorch, Scikit-Learn, XGBoost, and ONNX for real-time scoring and batch prediction.
  • Scale-to-Zero Pods - Scales inference pods down to zero replicas when idle and recreates them on demand using Knative Serving.
  • Model Serving Platforms - Defines model serving workloads as Kubernetes custom resources with automatic scaling, versioning, and traffic management.
  • Model Inference Deployments - Deploys trained models as scalable inference services on Kubernetes with full lifecycle management.
  • Standardized Inference Protocols - Exposes inference endpoints that comply with industry-standard protocols for client compatibility without custom adapters.
  • Weighted Traffic Splitting - Distributes incoming requests among multiple targets according to a weighted allocation.
  • Revision Traffic Splits - Distributes request traffic across model revisions using weighted routing for gradual rollouts and A/B testing.
  • Canary Version Rollouts - Rolls out new model versions gradually and runs A/B tests to validate performance before full release.
  • Inference Graph Routing - Routes requests through a directed acyclic graph of model nodes supporting conditional branching and parallel execution.
  • Model Pipeline Orchestration - Routes inference requests through a directed graph of multiple models, chaining outputs as inputs to downstream steps.
  • Inference Protocol Standards - Uses a standardized request/response API for inference across all supported model frameworks.
  • OpenAI-Compatible Servers - Accepts requests at standard chat completion, completion, embeddings, and scoring endpoints for LLM applications.
  • Pre- and Post-Inference Pipelines - Runs custom transformation pipelines for feature engineering and post-processing on predictions.
  • Model Lineage Trackers - Manages model versioning and lineage through custom resources, tracking metadata and lifecycle history.
  • Multi-GPU Parallelism Strategies - Distributes model layers or replicates models across multiple GPUs using tensor, data, or expert parallelism.
  • Pipeline Automation Layers - Integrates model serving into Kubeflow-based ML workflows for end-to-end pipeline automation.
  • Multi-Node Inference Scaling - Distributes model inference across multiple worker nodes for parallel execution of large or sharded models.
  • In-Memory Model Caching - Keeps frequently used models in memory to reduce load times and improve inference latency.
  • Inference Traffic Routing - Creates gateways and HTTPRoutes to direct external traffic to LLM schedulers for controlled model access.
  • Inference Pipeline Orchestrators - Orchestrates complex inference workflows by chaining multiple models into ensembles, pipelines, and conditional routing graphs.
  • Kubernetes Deployments - Runs PyTorch models through the Triton Inference Server for high-performance serving in Kubernetes.
  • Activation and KV Cache Offloaders - Offloads key-value cache from GPU memory to CPU or disk to reduce memory usage during long-context inference.
  • KV-Cache-Aware Request Routing - Directs inference requests to GPUs with relevant cached context for reduced latency.
  • Multi-Source Model Loaders - Downloads model artifacts from cloud storage, HTTP, Git, PVCs, or Hugging Face into a local serving directory.
  • Custom Resource GPU Scheduling - Deploys large language models using a dedicated custom resource that supports prefix-aware routing, disaggregated serving, and fine-grained GPU scheduling.
  • Prefix-Aware and Disaggregated Deployments - Deploys large language models with prefix-aware routing and disaggregated serving for advanced generative inference capabilities.
  • ONNX Model Servers - Loads and serves models in the ONNX format for inference using a standardized runtime.
  • Model Format Translators - Receives requests in one LLM provider's format and transforms them to match the expected format of a different upstream provider.
  • Inference Throughput Optimizations - Configures autoscaling thresholds and batch settings to improve throughput and cost efficiency for fixed-size prediction workloads.
  • Ensemble Aggregations - Scores several models independently and merges their outputs into a single response using a configurable strategy.
  • Canary and Pipeline Rollouts - Rolls out models using canary deployments, inference pipelines, and ensembles to manage traffic and risk.
  • Custom Inference Containers - Packages arbitrary inference code into container images and runs them as scalable Kubernetes services.
  • Scikit-Learn Deployments - Packages trained Scikit-Learn estimators into containers and serves them as scalable inference endpoints.
  • XGBoost Deployments - Packages and serves trained XGBoost models as scalable inference endpoints on Kubernetes.
  • Versioned Model Endpoints - Hosts several versions of the same model simultaneously and routes requests to the correct version.
  • Custom Runtime Templates - Defines reusable pod templates for serving custom model formats without modifying controller code.
  • Format-Based Runtime Selection - Automatically picks a ServingRuntime that supports the declared model format and version when no runtime is explicitly named.
  • Named Runtime Selections - Specifies a named serving runtime in deployment configuration to control which runtime serves the model.
  • Pre- and Post-Inference Transformations - Transforms input data and model output through pluggable pipelines for feature engineering and data preparation.
  • LoRA Adapter Loaders - Loads task-specific LoRA adapters alongside base models for efficient multi-tenant fine-tuning and serving.
  • Prefill-Decode Disaggregation - Separates the prefill and decode phases of LLM inference into distinct workloads for optimized resource utilization.
  • Sequential Step Orchestrators - Runs a series of model steps one after another, passing data conditionally between them to form multi-stage pipelines.
  • Parallel Step Executions - Executes multiple model steps concurrently and combines their individual responses into a single aggregated result.
  • Kubernetes Deployments - Deploys trained TensorFlow models as scalable inference services on Kubernetes using the standard SavedModel format.
  • Cached Model Deployments - Provides a mechanism to deploy inference services using models pre-cached on local node storage for faster startup.
  • Model Drift and Outlier Detection - Detects data drift, outliers, and collects inference metrics with Prometheus, OpenTelemetry, and payload logging for deployed models.
  • Inference Graph Chains - Connects router nodes so the output of one becomes the input of another, enabling complex model pipelines.
  • Model Caches - Stores frequently used models in a cache to reduce load times and improve inference latency.
  • LLM Model Caches - Pre-downloads large language models to Kubernetes node local NVMe volumes, reducing InferenceService startup time.
  • Model Cache Node Groups - Stores model data on individual cluster nodes to reduce latency and improve inference performance.
  • Cache Resource Specifications - KServe creates a LocalModelCache resource that defines the source model storage URI and target node groups to pre-download models to local volumes.
  • Inference Batching - Groups multiple prediction requests into a single batch to improve throughput on GPU and CPU runtimes.
  • MLflow Deployments - Deploys MLflow-packaged models as scalable inference services on Kubernetes.
  • Inference Routing - Routes model inference traffic through a service mesh to enforce traffic policies, security rules, and observability.
  • Control and Data Plane Separation - Separates management of service lifecycle from request execution for independent scaling and reliability.
  • GPU Resource Allocators - Allocates GPU resources, higher memory, and longer timeouts to meet the computational demands of content generation.
  • Model Inference Deployments - Deploys a model as a routable inference service on OpenShift using a declarative Kubernetes custom resource.
  • Local Weight Caches - Caches large model files on local nodes to cut startup time from minutes to under a minute.
  • Knative or Standard Deployments - Selects either standard Kubernetes resources or Knative Serving to match production stability or scale-to-zero needs.
  • Token-Based Rate Limiters - Limits token consumption per user per model using a global rate limit policy, returning a 429 error when exceeded.
  • Resource-Metric-Based Auto Scaling - Adjusts replica counts automatically using CPU, memory, custom metrics, or external event-driven triggers.
  • Gateway API Integrations - Manages traffic ingress and egress using the Kubernetes Gateway API for streaming and long-lived connections.
  • gRPC Event Communication - Exchanges inference data over gRPC for low-latency, high-throughput communication between clients and models.
  • Token Streaming - Sends output token by token using Server-Sent Events, allowing clients to see partial results before generation completes.
  • Gateway API Traffic Configurations - Uses Gateway API resources to enable traffic splitting, canary deployments, and protocol-specific routing.
  • Model Inference Conditionals - Evaluates input conditions to route each request to the correct model step in a multi-model inference graph.
  • Inference Endpoint Authentication - Restricts access to inference endpoints with authentication and authorization controls.
  • Inference Endpoint Access Controls - Restricts access to the inference service using RBAC and authenticates requests before routing them to the model.
  • Storage Backends - Configures how models are stored and retrieved from cloud storage, local volumes, and caches to reduce latency.
  • AI Service Access Policies - Validates API keys and enforces fine-grained backend security policies to control access to LLM/AI services.
  • Storage Credential Providers - Supports service account credentials, access keys, and custom certificates for secure model retrieval.
  • API Key Authentications - Verifies inference requests using simple API key tokens for lightweight access control.
  • OAuth2 Providers - Verifies inference requests using OAuth2 or OpenID Connect tokens for secure access control.
  • Model Performance Monitoring - Monitors deployed models with payload logging, outlier detection, adversarial detection, and drift detection.
  • Request Tracing - Captures request and response payloads and traces each prediction end-to-end for debugging and audit.
  • Inference Data Transformers - Pre-processes and post-processes inference data through an optional transformer component.
  • Kubernetes Deployments - Runs inference services as standard Kubernetes Deployments for maximum control and enterprise compatibility.
  • OpenAI-Compatible API Servers - Exposes an API compatible with OpenAI endpoints for chat completions, embeddings, and text generation from deployed models.
  • Machine Learning Operations - Standardized serverless inference platform for deploying and serving machine learning models on Kubernetes.
  • Machine Learning Platforms - Serverless platform for deploying machine learning models on Kubernetes.
  • MLOps Platforms - Provides serverless ML inference on Kubernetes.
  • Model Serving & Deployment - Serves predictive and generative models on Kubernetes.
5,576 stars·1,530 forks·Go·Apache-2.0·10 vues

Historique des stars

Graphique de l'historique des stars pour kserve/kserveGraphique de l'historique des stars pour kserve/kserve

Alternatives open source à Kserve

Projets open source similaires, classés selon le nombre de fonctionnalités partagées avec Kserve.
  • kubeflow/kfservingAvatar de kubeflow

    kubeflow/kfserving

    5,576Voir sur GitHub↗

    KServe is an open platform for deploying and serving generative and predictive AI models on Kubernetes. It defines inference services as custom resources with declarative YAML specifications, enabling a Kubernetes-native approach to model deployment and lifecycle management. The platform leverages Knative-based serverless scaling for automatic scale-to-zero and revision management, and supports a pluggable serving runtime architecture that maps model formats to containerized execution environments. KServe distinguishes itself through model-aware autoscaling that scales replicas based on token

    Go
    Voir sur GitHub↗5,576
  • seldonio/seldon-coreAvatar de SeldonIO

    SeldonIO/seldon-core

    4,752Voir sur GitHub↗

    Seldon Core is a Kubernetes-based machine learning model server and MLOps inference framework. It functions as a multi-model serving engine and pipeline orchestrator, packaging models as scalable microservices that are exposed via standardized REST and gRPC APIs. The project distinguishes itself through graph-based inference pipelines that chain models and data transformers into sequential workflows. It optimizes hardware utilization via multi-model shared serving and dynamic memory overcommit strategies, while supporting production experimentation through weighted traffic routing, A/B testin

    Goaiopsdeploymentkubernetes
    Voir sur GitHub↗4,752
  • dusty-nv/jetson-inferenceAvatar de dusty-nv

    dusty-nv/jetson-inference

    8,734Voir sur GitHub↗

    jetson-inference is a set of libraries and tools for executing optimized deep learning models on embedded GPU hardware. Its primary purpose is to enable real-time computer vision and AI inference at the edge with low latency and high throughput. The project distinguishes itself through high-performance streaming analytics and the ability to execute concurrent AI pipelines on auto-grade silicon. It provides specialized support for multi-sensor stream processing, utilizing zero-copy data transport to load camera frames directly into GPU memory. The codebase covers a broad surface of capabiliti

    C++caffecomputer-visiondeep-learning
    Voir sur GitHub↗8,734
  • pytorch/serveAvatar de pytorch

    pytorch/serve

    4,354Voir sur GitHub↗

    This project is a PyTorch model serving framework designed to deploy and scale machine learning models in production via scalable network endpoints. It functions as a high-performance inference server, optimizer, and model lifecycle manager that handles model loading, request batching, and hardware acceleration. The system distinguishes itself through advanced orchestration and optimization capabilities, such as chaining multiple models into sequential workflows using execution graphs and employing dynamic batching to improve throughput and latency. It provides specialized support for generat

    Java
    Voir sur GitHub↗4,354
Voir les 30 alternatives à Kserve→

Questions fréquentes

Que fait kserve/kserve ?

KServe is a Kubernetes-native platform for deploying and serving machine learning models as scalable inference services. It supports both generative AI models, including large language models, and traditional predictive models from frameworks such as TensorFlow, PyTorch, Scikit-Learn, XGBoost, and ONNX. The platform manages the full lifecycle of model deployments, including revision tracking, canary rollouts, A/B testing, and automatic rollbacks, and provides serverless…

Quelles sont les fonctionnalités principales de kserve/kserve ?

Les fonctionnalités principales de kserve/kserve sont : Kubernetes ML Platforms, OpenAI-Compatible APIs, KServe, Model Request Routing, Multi-Model Pipeline Routings, AI Request Routing, Generative AI Model Serving, GPU-Accelerated Inference.

Quelles sont les alternatives open-source à kserve/kserve ?

Les alternatives open-source à kserve/kserve incluent : kubeflow/kfserving — KServe is an open platform for deploying and serving generative and predictive AI models on Kubernetes. It defines… seldonio/seldon-core — Seldon Core is a Kubernetes-based machine learning model server and MLOps inference framework. It functions as a… dusty-nv/jetson-inference — jetson-inference is a set of libraries and tools for executing optimized deep learning models on embedded GPU… pytorch/serve — This project is a PyTorch model serving framework designed to deploy and scale machine learning models in production… weaveworks/flagger — Flagger is a Kubernetes operator designed to automate the lifecycle of application deployments through progressive… knative/serving — Kubernetes-based, scale-to-zero, request-driven compute.