awesome-repositories.com
المدونة
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
bentoml avatar

bentoml/OpenLLM

0
View on GitHub↗
12,115 نجوم·798 تفرعات·Python·apache-2.0·5 مشاهداتbentoml.com↗

OpenLLM

OpenLLM is a framework for deploying, managing, and scaling open-source large language models

Features

  • Serving Frameworks - Provides a platform for deploying, managing, and scaling open-source large language models as standardized API endpoints for production applications.
  • Large Language Models - Deploys and hosts open-source language models as standardized API endpoints to integrate artificial intelligence into production applications.
  • Model Inference Servers - Provides a system for hosting machine learning models with automated infrastructure provisioning, health monitoring, and elastic resource scaling.
  • Model Gateways - Provides a centralized interface for routing requests across multiple language models to simplify performance optimization and cost tracking.
  • AI Infrastructure Managers - Automates the deployment, scaling, and monitoring of machine learning models across cloud environments to ensure reliable performance.
  • AI Workflow Orchestrators - Connects multiple language models to build complex automated systems like retrieval-augmented generation pipelines.
  • LLM Application Orchestration - Chains multiple language models together to build complex automated pipelines and multi-step reasoning tasks.
  • Model Serving APIs - Exposes large language models through standard interface specifications to ensure seamless compatibility with development tools.
  • Open Models - Hosts popular open-source large language models as ready-to-use endpoints with pre-configured settings.
  • Retrieval Augmented Generation Pipelines - Connects multiple language models and data sources to build complex automated reasoning systems and advanced information retrieval workflows.
  • Cloud Deployment - Automates the transfer of hosted language models to managed cloud infrastructure to ensure scalable inference and reliable performance.
  • Container Orchestrators - Deploys model services as isolated containers that scale automatically based on incoming request volume and resource utilization metrics.
  • Model Registries - Maintains a searchable catalog of model definitions that allows for the hot-swapping and versioning of inference services at runtime.
  • Local Model Inference Servers - Runs large language models as local servers that provide standard-compliant APIs for easy integration.
  • Reasoning Pipelines - Connects multiple model endpoints into sequential execution chains to facilitate complex tasks like retrieval-augmented generation and multi-step reasoning.
  • Artificial Intelligence - API endpoint provider for running open-source LLMs.
  • Inference and Serving - Tool for serving models as OpenAI-compatible endpoints.
  • Inference Engines - Tool for serving open-source models as API endpoints.
  • Infrastructure and Gateways - Platform for operating LLMs in production.
  • Large Language Models - Platform for serving, deploying, and monitoring LLMs in production.
  • MLOps Platforms - Operates and deploys large language models in production.
  • Model Deployment and Platforms - Open platform for operating large language models in production.
  • Model Serving - Platform for operating and deploying large language models.
  • Model Serving and Inference - Tool for serving open-source LLMs as API endpoints.
  • Model Serving & Deployment - Runs open-source LLMs as OpenAI-compatible APIs.
  • Inference Frameworks - Deployment framework supporting multiple adapters and LangChain.
  • Private Cloud Deployments - Automates the setup of cloud-based inference environments with autoscaling and monitoring to support both fully-managed and private infrastructure.
  • Inference Scaling Services - Adjusts compute capacity through elastic auto-scaling and cross-region orchestration to optimize performance for production AI workloads.
  • Model Deployment Management - Controls versioning, rollbacks, and traffic shifting strategies like canary testing to ensure safe and reliable updates for production services.
  • Custom Model Architectures - Packages and hosts fine-tuned or custom model architectures using a standardized serving interface.
  • Model Configuration - Uses structured metadata and engine configurations to package and deploy new open-source language models as standardized services.
  • Model Packaging - Uses structured metadata files to define model configurations and dependencies for consistent deployment across diverse infrastructure environments.
  • LLM Performance Monitoring - Tracks system performance, compute utilization, and model-specific metrics for production AI services.
  • Third-party API Clients - Exposes model functionality through common interface specifications to ensure compatibility with existing development tools and third-party applications.
  • Chat Interfaces - Provides a web-based environment for interacting with hosted models and managing concurrent conversation threads.
  • Model Repositories - Connects external version control repositories containing model definitions to extend the local library with custom collections.
  • Model Management - Maintains a searchable registry of available language models and supports custom repositories to expand the collection of runnable software.

سجل النجوم

مخطط تاريخ النجوم لـ bentoml/openllmمخطط تاريخ النجوم لـ bentoml/openllm

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

الأسئلة الشائعة

ما هي وظيفة bentoml/openllm؟

OpenLLM is a framework for deploying, managing, and scaling open-source large language models

ما هي الميزات الرئيسية لـ bentoml/openllm؟

الميزات الرئيسية لـ bentoml/openllm هي: Serving Frameworks, Large Language Models, Model Inference Servers, Model Gateways, AI Infrastructure Managers, AI Workflow Orchestrators, LLM Application Orchestration, Model Serving APIs.

ما هي البدائل مفتوحة المصدر لـ bentoml/openllm؟

تشمل البدائل مفتوحة المصدر لـ bentoml/openllm: sgl-project/sglang — Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It… huggingface/text-generation-inference — Text Generation Inference is a production-ready engine designed for the deployment and serving of large language… vllm-project/vllm — vLLM is a high-throughput inference engine designed for the efficient serving and execution of large language models.… berriai/litellm — LiteLLM is a unified gateway and proxy server designed to centralize access to over one hundred language model… open-webui/open-webui — Open WebUI is a self-hosted, web-based platform designed for interacting with local and remote artificial intelligence… mudler/localai — LocalAI is a self-hosted inference server that enables the execution of machine learning models directly on local…

بدائل مفتوحة المصدر لـ OpenLLM

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع OpenLLM.
  • sgl-project/sglangالصورة الرمزية لـ sgl-project

    sgl-project/sglang

    29,079عرض على GitHub↗

    Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It provides a programmable interface for orchestrating complex generation workflows, enabling developers to coordinate multi-turn dialogues, tool invocations, and reasoning chains through a domain-specific language. The platform is built to support production-scale deployments, offering an OpenAI-compatible API that allows for integration with existing application ecosystems. The system distinguishes itself through a disaggregated architecture that separates compute-intensive pr

    Pythonattentionblackwellcuda
    عرض على GitHub↗29,079
  • huggingface/text-generation-inferenceالصورة الرمزية لـ huggingface

    huggingface/text-generation-inference

    10,775عرض على GitHub↗

    Text Generation Inference is a production-ready engine designed for the deployment and serving of large language models. It functions as a containerized runtime environment that manages model execution, scales across distributed hardware, and provides high-performance inference capabilities for demanding production environments. The project distinguishes itself through advanced optimization techniques, including continuous batching to maximize hardware utilization and tensor parallelism to shard large models across multiple accelerator cards. It supports efficient inference through custom com

    Pythonbloomdeep-learningfalcon
    عرض على GitHub↗10,775
  • vllm-project/vllmالصورة الرمزية لـ vllm-project

    vllm-project/vllm

    83,048عرض على GitHub↗

    vLLM is a high-throughput inference engine designed for the efficient serving and execution of large language models. It functions as a production-ready distributed model server, providing standard API protocols for online serving while also supporting offline batch processing. The system is built to maximize token generation speed and memory efficiency, enabling both large-scale cloud deployments and local execution on personal hardware. The project distinguishes itself through advanced memory management and request scheduling techniques, most notably its use of non-contiguous key-value cach

    Pythonamdblackwellcuda
    عرض على GitHub↗83,048
  • berriai/litellmالصورة الرمزية لـ BerriAI

    BerriAI/litellm

    50,579عرض على GitHub↗

    LiteLLM is a unified gateway and proxy server designed to centralize access to over one hundred language model providers. It provides a standardized API interface that abstracts vendor-specific schemas, allowing developers to interact with diverse models through a single, consistent format. By acting as a central traffic management layer, it enables organizations to route, secure, and govern model interactions across multiple deployments. The platform distinguishes itself through its policy-driven architecture, which uses configuration-based routing to manage traffic distribution, load balanc

    Pythonai-gatewayanthropicazure-openai
    عرض على GitHub↗50,579
  • عرض جميع البدائل الـ 30 لـ OpenLLM→