3 रिपॉजिटरी
Specifies a content-level classification and semantic routing framework for AI inference systems as an IETF protocol.
Distinct from Routing Protocols: No candidate covers an IETF protocol specification for AI inference routing; closest are network routing protocols.
Explore 3 awesome GitHub repositories matching artificial intelligence & ml · Inference Routing Protocols. Refine with filters or upvote what's useful.
KServe is a Kubernetes-native platform for deploying and serving machine learning models as scalable inference services. It supports both generative AI models, including large language models, and traditional predictive models from frameworks such as TensorFlow, PyTorch, Scikit-Learn, XGBoost, and ONNX. The platform manages the full lifecycle of model deployments, including revision tracking, canary rollouts, A/B testing, and automatic rollbacks, and provides serverless scale-to-zero capabilities for cost-efficient resource management. KServe distinguishes itself through a standardized infere
Standardizes inference requests and responses across REST and gRPC with health checking and metadata endpoints.
Seldon Core एक Kubernetes-आधारित मशीन लर्निंग मॉडल सर्वर और MLOps इन्फरेंस फ्रेमवर्क है। यह एक मल्टी-मॉडल सर्विंग इंजन और पाइपलाइन ऑर्केस्ट्रेटर के रूप में कार्य करता है, जो मॉडल्स को स्केलेबल माइक्रोसर्विसेज के रूप में पैकेज करता है जिन्हें स्टैंडर्ड REST और gRPC API के माध्यम से एक्सपोज़ किया जाता है। यह प्रोजेक्ट ग्राफ-आधारित इन्फरेंस पाइपलाइन्स के माध्यम से अलग है जो मॉडल्स और डेटा ट्रांसफॉर्मर्स को अनुक्रमिक वर्कफ़्लो में जोड़ते हैं। यह मल्टी-मॉडल शेयर्ड सर्विंग और डायनामिक मेमोरी ओवरकमिट रणनीतियों के माध्यम से हार्डवेयर उपयोग को ऑप्टिमाइज़ करता है, जबकि वेटेड ट्रैफिक रूटिंग, A/B टेस्टिंग और शैडो डिप्लॉयमेंट के माध्यम से प्रोडक्शन एक्सपेरिमेंटेशन का समर्थन करता है। यह फ्रेमवर्क डिमांड-आधारित ऑटोस्केलिंग, मैसेज बसों के माध्यम से एसिंक्रोनस रिक्वेस्ट प्रोसेसिंग, और डेटा ड्रिफ्ट, आउटलेयर्स और प्रेडिक्शन एक्सप्लेनबिलिटी के लिए व्यापक मॉनिटरिंग सहित MLOps क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। यह मॉडल रनटाइम कॉन्फ़िगरेशन के लिए इंफ्रास्ट्रक्चर मैनेजमेंट और कंट्रोल और डेटा प्लेन्स पर TLS एन्क्रिप्शन का उपयोग करके सुरक्षित संचार भी प्रदान करता है।
Implements standardized REST and gRPC protocols for consistent request and response handling across different model runtimes.
Specifies a content-level classification and semantic routing framework as an IETF protocol.