8 Repos
Connectors for executing models on managed cloud inference services.
Distinguishing note: Focuses on specific cloud inference platforms, distinct from general API providers.
Explore 8 awesome GitHub repositories matching artificial intelligence & ml · Inference Endpoint Integrations. Refine with filters or upvote what's useful.
LiteLLM is a unified gateway and proxy server designed to centralize access to over one hundred language model providers. It provides a standardized API interface that abstracts vendor-specific schemas, allowing developers to interact with diverse models through a single, consistent format. By acting as a central traffic management layer, it enables organizations to route, secure, and govern model interactions across multiple deployments. The platform distinguishes itself through its policy-driven architecture, which uses configuration-based routing to manage traffic distribution, load balanc
Executes chat completions and multimodal analysis on real-time inference endpoints.
Kotaemon is an orchestration framework designed for building modular, agentic workflows that integrate document processing, retrieval-augmented generation, and multi-step reasoning. It provides a comprehensive platform for developing document-based question answering systems, allowing users to chain language models, prompt templates, and external tools into complex, automated pipelines. The system distinguishes itself through a highly modular architecture that emphasizes component-based composition and schema-driven data exchange. It supports autonomous agents capable of decomposing complex q
Connects to compatible API endpoints to utilize custom or hosted language models for document processing.
Pipecat is a framework and software development kit for building real-time multimodal AI agents and speech-to-speech systems. It utilizes a frame-based data pipeline to route audio, video, and text through a modular sequence of processors, enabling the orchestration of low-latency conversational AI. The project is distinguished by its ability to coordinate complex multimodal services, including speech-to-text, language models, and text-to-speech, within a single pipeline. It features semantic voice activity detection for natural turn-taking, state-machine conversation flows for dialogue manag
Connects to bidirectional-stream endpoints on AWS SageMaker for real-time speech recognition.
Moondream is a small-scale vision language model designed to reason across images to generate captions and answer natural language questions. It functions as an edge-optimized system capable of performing visual question answering, image captioning, and object detection. The project distinguishes itself through a lightweight architecture designed for local inference on embedded devices, workstations, and air-gapped hardware. It supports the execution of models on local GPUs and Apple Silicon to ensure data privacy and low latency. The system's capabilities include identifying precise object
Provides autoscaled inference endpoints for fine-tuned models accessible via a standard SDK.
This project is an AI-powered IDE extension and LLM coding assistant that provides a conversational interface for generating, refactoring, and debugging code. It functions as an AI agent framework and a Model Context Protocol client, connecting AI models to external data sources and tools to automate complex development tasks. The system is distinguished by its use of autonomous AI agents capable of multi-step task execution, including the ability to read files, modify code, and run terminal commands iteratively. It supports recursive agent orchestration through subagent delegation and employ
Provisions cloud endpoints for deploying models with trained adapters to enable web API interactions.
Agones is a Kubernetes game server orchestrator designed for hosting, scaling, and managing dedicated multiplayer game servers. It extends the Kubernetes control plane using custom resource definitions to define game server and fleet objects, utilizing a dedicated fleet manager to maintain pools of warm server instances. The system provides a game server SDK and language-specific client libraries that allow server processes to signal readiness, health, and shutdown states directly to the controller. It distinguishes itself through specialized scaling logic, including the use of WebAssembly mo
Provides the capability to connect game server instances to external AI inference endpoints for real-time interactions.
Dieses Projekt ist ein Software Development Kit (SDK) und Cluster-Management-Tool für PHP. Es dient als SDK für Volltextsuche und Vektor-Suchschnittstelle, wodurch Anwendungen lexikalische, Fuzzy- und semantische Suchen auf indizierten Daten durchführen können. Die Bibliothek implementiert einen PSR-7-HTTP-Client, um die Kompatibilität zwischen verschiedenen Umgebungen durch standardisierte Messaging-Schnittstellen zu gewährleisten. Sie bietet eine spezialisierte Schnittstelle zum Abrufen von Embeddings und zur Durchführung semantischer Retrieval-Workflows unter Verwendung von Vektordaten. Der Funktionsumfang deckt eine breite Palette administrativer und operativer Aufgaben ab, einschließlich der Verwaltung von Suchindizes, der Überwachung des Cluster-Status und der Verwaltung von Dokumentlebenszyklen. Es unterstützt diverse Abfragemethoden wie SQL, EQL und ES|QL sowie Datenaggregation und Geodatenanalyse. Zusätzlich bietet es Tools für Machine-Learning-Orchestrierung, Anomalieerkennung sowie Identitäts- und Zugriffsmanagement.
Supports the creation and configuration of endpoints for running models from external or local providers.
The Hugging Face Hub Python client is a library that provides programmatic access to the Hugging Face Hub, a centralized platform for hosting and collaborating on machine learning models, datasets, and demo applications. It serves as the primary SDK for interacting with the Hub's API, enabling users to download and upload models and datasets, manage repositories, authenticate via tokens or OAuth, and run inference on hosted models through a unified interface. The client distinguishes itself through a comprehensive set of capabilities that go beyond basic file transfer. It includes a CLI exten
Creates, configures, and manages remote inference endpoints for scalable model serving.