awesome-repositories.com
Blog
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectDespreCum realizăm clasamentulPresăServer MCP
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
bentoml avatar

bentoml/OpenLLM

0
View on GitHub↗
12,115 stele·798 fork-uri·Python·apache-2.0·5 vizualizăribentoml.com↗

OpenLLM

OpenLLM is a framework for deploying, managing, and scaling open-source large language models

Features

  • Serving Frameworks - Provides a platform for deploying, managing, and scaling open-source large language models as standardized API endpoints for production applications.
  • Large Language Models - Deploys and hosts open-source language models as standardized API endpoints to integrate artificial intelligence into production applications.
  • Model Inference Servers - Provides a system for hosting machine learning models with automated infrastructure provisioning, health monitoring, and elastic resource scaling.
  • Model Gateways - Provides a centralized interface for routing requests across multiple language models to simplify performance optimization and cost tracking.
  • AI Infrastructure Managers - Automates the deployment, scaling, and monitoring of machine learning models across cloud environments to ensure reliable performance.
  • AI Workflow Orchestrators - Connects multiple language models to build complex automated systems like retrieval-augmented generation pipelines.
  • LLM Application Orchestration - Chains multiple language models together to build complex automated pipelines and multi-step reasoning tasks.
  • Model Serving APIs - Exposes large language models through standard interface specifications to ensure seamless compatibility with development tools.
  • Open Models - Hosts popular open-source large language models as ready-to-use endpoints with pre-configured settings.
  • Retrieval Augmented Generation Pipelines - Connects multiple language models and data sources to build complex automated reasoning systems and advanced information retrieval workflows.
  • Cloud Deployment - Automates the transfer of hosted language models to managed cloud infrastructure to ensure scalable inference and reliable performance.
  • Container Orchestrators - Deploys model services as isolated containers that scale automatically based on incoming request volume and resource utilization metrics.
  • Model Registries - Maintains a searchable catalog of model definitions that allows for the hot-swapping and versioning of inference services at runtime.
  • Local Model Inference Servers - Runs large language models as local servers that provide standard-compliant APIs for easy integration.
  • Reasoning Pipelines - Connects multiple model endpoints into sequential execution chains to facilitate complex tasks like retrieval-augmented generation and multi-step reasoning.
  • Artificial Intelligence - API endpoint provider for running open-source LLMs.
  • Inference and Serving - Tool for serving models as OpenAI-compatible endpoints.
  • Inference Engines - Tool for serving open-source models as API endpoints.
  • Infrastructure and Gateways - Platform for operating LLMs in production.
  • Large Language Models - Platform for serving, deploying, and monitoring LLMs in production.
  • MLOps Platforms - Operates and deploys large language models in production.
  • Model Deployment and Platforms - Open platform for operating large language models in production.
  • Model Serving - Platform for operating and deploying large language models.
  • Model Serving and Inference - Tool for serving open-source LLMs as API endpoints.
  • Model Serving & Deployment - Runs open-source LLMs as OpenAI-compatible APIs.
  • Inference Frameworks - Deployment framework supporting multiple adapters and LangChain.
  • Private Cloud Deployments - Automates the setup of cloud-based inference environments with autoscaling and monitoring to support both fully-managed and private infrastructure.
  • Inference Scaling Services - Adjusts compute capacity through elastic auto-scaling and cross-region orchestration to optimize performance for production AI workloads.
  • Model Deployment Management - Controls versioning, rollbacks, and traffic shifting strategies like canary testing to ensure safe and reliable updates for production services.
  • Custom Model Architectures - Packages and hosts fine-tuned or custom model architectures using a standardized serving interface.
  • Model Configuration - Uses structured metadata and engine configurations to package and deploy new open-source language models as standardized services.
  • Model Packaging - Uses structured metadata files to define model configurations and dependencies for consistent deployment across diverse infrastructure environments.
  • LLM Performance Monitoring - Tracks system performance, compute utilization, and model-specific metrics for production AI services.
  • Third-party API Clients - Exposes model functionality through common interface specifications to ensure compatibility with existing development tools and third-party applications.
  • Chat Interfaces - Provides a web-based environment for interacting with hosted models and managing concurrent conversation threads.
  • Model Repositories - Connects external version control repositories containing model definitions to extend the local library with custom collections.
  • Model Management - Maintains a searchable registry of available language models and supports custom repositories to expand the collection of runnable software.

Istoric stele

Graficul istoricului de stele pentru bentoml/openllmGraficul istoricului de stele pentru bentoml/openllm

Căutare AI

Explorează mai multe repository-uri excelente

Descrie ce ai nevoie în limbaj simplu — AI-ul sortează mii de proiecte open source selectate în funcție de relevanță.

Start searching with AI

Alternative open-source pentru OpenLLM

Proiecte open-source similare, clasificate după numărul de funcționalități comune cu OpenLLM.
  • sgl-project/sglangAvatar sgl-project

    sgl-project/sglang

    29,079Vezi pe GitHub↗

    Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It provides a programmable interface for orchestrating complex generation workflows, enabling developers to coordinate multi-turn dialogues, tool invocations, and reasoning chains through a domain-specific language. The platform is built to support production-scale deployments, offering an OpenAI-compatible API that allows for integration with existing application ecosystems. The system distinguishes itself through a disaggregated architecture that separates compute-intensive pr

    Pythonattentionblackwellcuda
    Vezi pe GitHub↗29,079
  • huggingface/text-generation-inferenceAvatar huggingface

    huggingface/text-generation-inference

    10,775Vezi pe GitHub↗

    Text Generation Inference is a production-ready engine designed for the deployment and serving of large language models. It functions as a containerized runtime environment that manages model execution, scales across distributed hardware, and provides high-performance inference capabilities for demanding production environments. The project distinguishes itself through advanced optimization techniques, including continuous batching to maximize hardware utilization and tensor parallelism to shard large models across multiple accelerator cards. It supports efficient inference through custom com

    Pythonbloomdeep-learningfalcon
    Vezi pe GitHub↗10,775
  • vllm-project/vllmAvatar vllm-project

    vllm-project/vllm

    83,048Vezi pe GitHub↗

    vLLM is a high-throughput inference engine designed for the efficient serving and execution of large language models. It functions as a production-ready distributed model server, providing standard API protocols for online serving while also supporting offline batch processing. The system is built to maximize token generation speed and memory efficiency, enabling both large-scale cloud deployments and local execution on personal hardware. The project distinguishes itself through advanced memory management and request scheduling techniques, most notably its use of non-contiguous key-value cach

    Pythonamdblackwellcuda
    Vezi pe GitHub↗83,048
  • berriai/litellmAvatar BerriAI

    BerriAI/litellm

    50,579Vezi pe GitHub↗

    LiteLLM is a unified gateway and proxy server designed to centralize access to over one hundred language model providers. It provides a standardized API interface that abstracts vendor-specific schemas, allowing developers to interact with diverse models through a single, consistent format. By acting as a central traffic management layer, it enables organizations to route, secure, and govern model interactions across multiple deployments. The platform distinguishes itself through its policy-driven architecture, which uses configuration-based routing to manage traffic distribution, load balanc

    Pythonai-gatewayanthropicazure-openai
    Vezi pe GitHub↗50,579
Vezi toate cele 30 alternative pentru OpenLLM→

Întrebări frecvente

Ce face bentoml/openllm?

OpenLLM is a framework for deploying, managing, and scaling open-source large language models

Care sunt principalele funcționalități ale bentoml/openllm?

Principalele funcționalități ale bentoml/openllm sunt: Serving Frameworks, Large Language Models, Model Inference Servers, Model Gateways, AI Infrastructure Managers, AI Workflow Orchestrators, LLM Application Orchestration, Model Serving APIs.

Care sunt câteva alternative open-source pentru bentoml/openllm?

Alternativele open-source pentru bentoml/openllm includ: sgl-project/sglang — Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It… huggingface/text-generation-inference — Text Generation Inference is a production-ready engine designed for the deployment and serving of large language… vllm-project/vllm — vLLM is a high-throughput inference engine designed for the efficient serving and execution of large language models.… berriai/litellm — LiteLLM is a unified gateway and proxy server designed to centralize access to over one hundred language model… open-webui/open-webui — Open WebUI is a self-hosted, web-based platform designed for interacting with local and remote artificial intelligence… mudler/localai — LocalAI is a self-hosted inference server that enables the execution of machine learning models directly on local…