awesome-repositories.com
Blog
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoAcerca deCómo clasificamosPrensaServidor MCP
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
OpenBMB avatar

OpenBMB/ToolBench

0
View on GitHub↗
5,672 estrellas·485 forks·Python·Apache-2.0·5 vistasopenbmb.github.io/ToolBench↗

ToolBench

ToolBench is an open platform for training, serving, and evaluating large language models that retrieve and call real-world APIs to complete user instructions. It provides an API-aware inference engine that selects relevant tools from a large corpus and generates sequences of tool calls to produce final answers, along with a custom API registration system that lets users add their own REST endpoints for the model to discover and invoke.

The platform includes a complete instruction-tuning pipeline for training models on curated tool-use data, a multi-tool execution engine that coordinates sequential and parallel API calls, and a preference-based quality assessment framework that uses a judge model to compare candidate action sequences. It also offers a tool-use evaluation benchmark that measures task completion rate and preference between different tool-calling strategies.

Beyond training and evaluation, ToolBench provides a web-based chat interface and REST endpoints for serving fine-tuned models with tool-augmented responses. The system supports API retrieval for open-domain queries, enabling models to handle tasks requiring any of thousands of tools, and allows users to register custom APIs with documentation and implementation code for model discovery during inference.

Features

  • API Retrieval-Augmented Generations - Retrieves relevant APIs from a large corpus before inference to ground model responses in real-world tools.
  • Tool-Use Training - Provides a complete instruction-tuning pipeline for training models on curated tool-use data.
  • Tool-Use Instruction Tuning - Includes a complete instruction-tuning pipeline for training models on curated tool-use data.
  • Instruction Tuning Pipelines - Provides a complete instruction-tuning pipeline for training models on curated tool-use data.
  • API-Aware Inference Engines - Provides an API-aware inference engine that selects relevant tools from a large corpus and generates tool-calling sequences.
  • Tool-Use Performance Evaluators - Measuring how well a model completes tasks using APIs, including task success rate and preference between candidate action sequences.
  • Tool-Augmented Serving Frameworks - Provides a web-based chat interface and REST endpoints for serving fine-tuned models with tool-augmented responses.
  • Tool Call Data Generators - Generates high-quality instruction-tuning examples that teach models to plan and execute API calls.
  • Tool-Using Model Inference - Provides an inference engine that retrieves APIs and generates tool-call sequences to answer user queries.
  • Tool-Use Serving Frameworks - Provides a web-based chat interface and REST endpoints for serving fine-tuned models with tool-augmented responses.
  • API-Use Fine-Tuning - Fine-tunes language models on curated instruction-solution path datasets to enable real-world REST API calling.
  • API Retrieval - Supports API retrieval for open-domain queries, enabling models to handle tasks requiring any of thousands of tools.
  • API Retrieval Systems - Selects relevant APIs from a large corpus before inference to ground model responses in real-world tools.
  • Custom API Registrations - Registers custom REST APIs with documentation and implementation so the model can discover and call them during inference.
  • API Registration Systems - Ships a custom API registration system that lets users add their own REST endpoints for model discovery and invocation.
  • API Registration Systems - Lets users register custom REST APIs with documentation and implementation for model discovery during inference.
  • Custom API Registrations - Registers user-defined REST APIs with documentation and implementation for model discovery during inference.
  • Multi-Tool Execution Engines - Coordinates sequential and parallel API calls across multiple tools to complete complex instructions.
  • Custom API Inferences - Allows registering user-defined APIs with documentation and implementation for use during inference.
  • Preference-Based Quality Assessments - Uses a ChatGPT-based judge to compare action sequences and determine preferred tool-use strategies.
  • Action-Sequence Evaluation Frameworks - Ships a judge-model framework that compares tool-call sequences to assess quality and preference.
  • Action Sequence Comparators - Ships a preference-based quality assessment framework that uses a judge model to compare candidate action sequences.
  • Tool-Use Evaluation Benchmarks - Offers a tool-use evaluation benchmark that measures task completion rate and preference between different tool-calling strategies.
  • Custom API Inferences - Extends the model's tool-use capability to user-defined APIs by providing their documentation and implementation code.
  • Tool-Using Model API Servings - Exposes a fine-tuned model behind a REST endpoint that streams back tool-augmented responses.
  • Tool-Use Accuracy Evaluators - Measures task completion rate for tool-using models, reporting the proportion of successfully completed instructions.
  • Tool-Use Quality Comparators - Implements a ChatGPT-based judge that compares action sequences to determine preferred tool-use strategies.
  • Web Chat Interfaces - Provides a web-based chat interface for interacting with a tool-using model.
  • Tool-Using Model Servings - Runs a fine-tuned model behind a web server that streams back tool-augmented responses.
  • Tool-Using Model Web UIs - Provides a chat interface where users interact with a tool-using model via a web UI.
  • Agent Tool Learning - Facilitates mastering real-world APIs for tool-augmented agents.
  • Language Model Development - Platform for training and evaluating LLMs for tool use.
  • Tool Optimization - Facilitating mastery of real-world APIs.
  • Tool Use And Integration - Facilitates training and evaluation for large-scale API tool use.

Historial de estrellas

Gráfico del historial de estrellas de openbmb/toolbenchGráfico del historial de estrellas de openbmb/toolbench

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Alternativas open-source a ToolBench

Proyectos open-source similares, clasificados según cuántas características comparten con ToolBench.
  • microsoft/jarvisAvatar de microsoft

    microsoft/JARVIS

    24,854Ver en GitHub↗

    JARVIS is a system for large language model task orchestration, deployment management, and automation benchmarking. It utilizes a task orchestrator to decompose complex requests into actionable steps and coordinates various expert models to synthesize final responses. The project includes an AI model deployment manager to handle the local deployment of expert models across different hardware scales. It further provides an AI workflow API consisting of web endpoints used to trigger automated task workflows and retrieve results from model selection stages. The framework incorporates an automat

    Python
    Ver en GitHub↗24,854
  • shizhl/confuciusAvatar de shizhl

    shizhl/Confucius

    48Ver en GitHub↗

    This project aims to teach large language models (LLMs) to use external tools, enabling them as agents to for practical task planning and tool calling. To achieve this, we propose the "Iterative Tool Learning from Introspection Feedback by Easy-to-Difficult Curriculum". Here is the brief…

    Python
    Ver en GitHub↗48
  • ailab-cvc/gpt4toolsAvatar de AILab-CVC

    AILab-CVC/GPT4Tools

    771Ver en GitHub↗

    GPT4Tools is an intelligent system that can automatically decide, control, and utilize different visual foundation models, allowing the user to interact with images during a conversation.

    Python
    Ver en GitHub↗771
  • evolvinglmms-lab/otterAvatar de EvolvingLMMs-Lab

    EvolvingLMMs-Lab/Otter

    3,331Ver en GitHub↗

    Otter is a framework and toolkit for the pretraining, fine-tuning, and evaluation of vision-language models. It provides a pipeline for training large language models to process high-resolution images and video frames, integrating visual encoders with textual token spaces. The system is designed for multi-visual input processing, allowing models to interpret multiple images or video sequences within a single prompt. It supports multi-round conversation management to maintain context across interactions for detailed scene comprehension and visual reasoning. The framework covers a full develop

    Pythonartificial-inteligencechatgptdeep-learning
    Ver en GitHub↗3,331
Ver las 30 alternativas a ToolBench→

Preguntas frecuentes

¿Qué hace openbmb/toolbench?

ToolBench is an open platform for training, serving, and evaluating large language models that retrieve and call real-world APIs to complete user instructions. It provides an API-aware inference engine that selects relevant tools from a large corpus and generates sequences of tool calls to produce final answers, along with a custom API registration system that lets users add their own REST endpoints for the model to discover and invoke.

¿Cuáles son las características principales de openbmb/toolbench?

Las características principales de openbmb/toolbench son: API Retrieval-Augmented Generations, Tool-Use Training, Tool-Use Instruction Tuning, Instruction Tuning Pipelines, API-Aware Inference Engines, Tool-Use Performance Evaluators, Tool-Augmented Serving Frameworks, Tool Call Data Generators.

¿Qué alternativas de código abierto existen para openbmb/toolbench?

Las alternativas de código abierto para openbmb/toolbench incluyen: microsoft/jarvis — JARVIS is a system for large language model task orchestration, deployment management, and automation benchmarking. It… shizhl/confucius — This project aims to teach large language models (LLMs) to use external tools, enabling them as agents to for… ailab-cvc/gpt4tools — GPT4Tools is an intelligent system that can automatically decide, control, and utilize different visual foundation… evolvinglmms-lab/otter — Otter is a framework and toolkit for the pretraining, fine-tuning, and evaluation of vision-language models. It… allenai/open-instruct — Open-Instruct is a distributed training and instruction tuning framework for large language models. It functions as a… gen-verse/openclaw-rl — OpenClaw-RL is a reinforcement learning framework for training large language model agents. It provides a system for…