awesome-repositories.com
Blog
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
OpenBMB avatar

OpenBMB/ToolBench

0
View on GitHub↗
5,672 Stars·485 Forks·Python·Apache-2.0·2 Aufrufeopenbmb.github.io/ToolBench↗

ToolBench

ToolBench is an open platform for training, serving, and evaluating large language models that retrieve and call real-world APIs to complete user instructions. It provides an API-aware inference engine that selects relevant tools from a large corpus and generates sequences of tool calls to produce final answers, along with a custom API registration system that lets users add their own REST endpoints for the model to discover and invoke.

The platform includes a complete instruction-tuning pipeline for training models on curated tool-use data, a multi-tool execution engine that coordinates sequential and parallel API calls, and a preference-based quality assessment framework that uses a judge model to compare candidate action sequences. It also offers a tool-use evaluation benchmark that measures task completion rate and preference between different tool-calling strategies.

Beyond training and evaluation, ToolBench provides a web-based chat interface and REST endpoints for serving fine-tuned models with tool-augmented responses. The system supports API retrieval for open-domain queries, enabling models to handle tasks requiring any of thousands of tools, and allows users to register custom APIs with documentation and implementation code for model discovery during inference.

Features

  • API Retrieval-Augmented Generations - Retrieves relevant APIs from a large corpus before inference to ground model responses in real-world tools.
  • Tool-Use Training - Provides a complete instruction-tuning pipeline for training models on curated tool-use data.
  • Tool-Use Instruction Tuning - Includes a complete instruction-tuning pipeline for training models on curated tool-use data.
  • Instruction Tuning Pipelines - Provides a complete instruction-tuning pipeline for training models on curated tool-use data.
  • API-Aware Inference Engines - Provides an API-aware inference engine that selects relevant tools from a large corpus and generates tool-calling sequences.
  • Tool-Use Performance Evaluators - Measuring how well a model completes tasks using APIs, including task success rate and preference between candidate action sequences.
  • Tool-Augmented Serving Frameworks - Provides a web-based chat interface and REST endpoints for serving fine-tuned models with tool-augmented responses.
  • Tool Call Data Generators - Generates high-quality instruction-tuning examples that teach models to plan and execute API calls.
  • Tool-Using Model Inference - Provides an inference engine that retrieves APIs and generates tool-call sequences to answer user queries.
  • Tool-Use Serving Frameworks - Provides a web-based chat interface and REST endpoints for serving fine-tuned models with tool-augmented responses.
  • API-Use Fine-Tuning - Fine-tunes language models on curated instruction-solution path datasets to enable real-world REST API calling.
  • API Retrieval - Supports API retrieval for open-domain queries, enabling models to handle tasks requiring any of thousands of tools.
  • API Retrieval Systems - Selects relevant APIs from a large corpus before inference to ground model responses in real-world tools.
  • Custom API Registrations - Registers custom REST APIs with documentation and implementation so the model can discover and call them during inference.
  • API Registration Systems - Ships a custom API registration system that lets users add their own REST endpoints for model discovery and invocation.
  • API Registration Systems - Lets users register custom REST APIs with documentation and implementation for model discovery during inference.
  • Custom API Registrations - Registers user-defined REST APIs with documentation and implementation for model discovery during inference.
  • Multi-Tool Execution Engines - Coordinates sequential and parallel API calls across multiple tools to complete complex instructions.
  • Custom API Inferences - Allows registering user-defined APIs with documentation and implementation for use during inference.
  • Preference-Based Quality Assessments - Uses a ChatGPT-based judge to compare action sequences and determine preferred tool-use strategies.
  • Action-Sequence Evaluation Frameworks - Ships a judge-model framework that compares tool-call sequences to assess quality and preference.
  • Action Sequence Comparators - Ships a preference-based quality assessment framework that uses a judge model to compare candidate action sequences.
  • Tool-Use Evaluation Benchmarks - Offers a tool-use evaluation benchmark that measures task completion rate and preference between different tool-calling strategies.
  • Custom API Inferences - Extends the model's tool-use capability to user-defined APIs by providing their documentation and implementation code.
  • Tool-Using Model API Servings - Exposes a fine-tuned model behind a REST endpoint that streams back tool-augmented responses.
  • Tool-Use Accuracy Evaluators - Measures task completion rate for tool-using models, reporting the proportion of successfully completed instructions.
  • Tool-Use Quality Comparators - Implements a ChatGPT-based judge that compares action sequences to determine preferred tool-use strategies.
  • Web Chat Interfaces - Provides a web-based chat interface for interacting with a tool-using model.
  • Tool-Using Model Servings - Runs a fine-tuned model behind a web server that streams back tool-augmented responses.
  • Tool-Using Model Web UIs - Provides a chat interface where users interact with a tool-using model via a web UI.
  • Agent Tool Learning - Facilitates mastering real-world APIs for tool-augmented agents.
  • Language Model Development - Platform for training and evaluating LLMs for tool use.
  • Tool Optimization - Facilitating mastery of real-world APIs.
  • Tool Use And Integration - Facilitates training and evaluation for large-scale API tool use.

Star-Verlauf

Star-Verlauf für openbmb/toolbenchStar-Verlauf für openbmb/toolbench

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Häufig gestellte Fragen

Was macht openbmb/toolbench?

ToolBench is an open platform for training, serving, and evaluating large language models that retrieve and call real-world APIs to complete user instructions. It provides an API-aware inference engine that selects relevant tools from a large corpus and generates sequences of tool calls to produce final answers, along with a custom API registration system that lets users add their own REST endpoints for the model to discover and invoke.

Was sind die Hauptfunktionen von openbmb/toolbench?

Die Hauptfunktionen von openbmb/toolbench sind: API Retrieval-Augmented Generations, Tool-Use Training, Tool-Use Instruction Tuning, Instruction Tuning Pipelines, API-Aware Inference Engines, Tool-Use Performance Evaluators, Tool-Augmented Serving Frameworks, Tool Call Data Generators.

Welche Open-Source-Alternativen gibt es zu openbmb/toolbench?

Open-Source-Alternativen zu openbmb/toolbench sind unter anderem: microsoft/jarvis — JARVIS is a system for large language model task orchestration, deployment management, and automation benchmarking. It… shizhl/confucius — This project aims to teach large language models (LLMs) to use external tools, enabling them as agents to for… ailab-cvc/gpt4tools — GPT4Tools is an intelligent system that can automatically decide, control, and utilize different visual foundation… evolvinglmms-lab/otter — Otter is a framework and toolkit for the pretraining, fine-tuning, and evaluation of vision-language models. It… allenai/open-instruct — Open-Instruct is a distributed training and instruction tuning framework for large language models. It functions as a… gen-verse/openclaw-rl — OpenClaw-RL is a reinforcement learning framework for training large language model agents. It provides a system for…

Open-Source-Alternativen zu ToolBench

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit ToolBench.
  • microsoft/jarvisAvatar von microsoft

    microsoft/JARVIS

    24,854Auf GitHub ansehen↗

    JARVIS is a system for large language model task orchestration, deployment management, and automation benchmarking. It utilizes a task orchestrator to decompose complex requests into actionable steps and coordinates various expert models to synthesize final responses. The project includes an AI model deployment manager to handle the local deployment of expert models across different hardware scales. It further provides an AI workflow API consisting of web endpoints used to trigger automated task workflows and retrieve results from model selection stages. The framework incorporates an automat

    Python
    Auf GitHub ansehen↗24,854
  • shizhl/confuciusAvatar von shizhl

    shizhl/Confucius

    48Auf GitHub ansehen↗

    This project aims to teach large language models (LLMs) to use external tools, enabling them as agents to for practical task planning and tool calling. To achieve this, we propose the "Iterative Tool Learning from Introspection Feedback by Easy-to-Difficult Curriculum". Here is the brief…

    Python
    Auf GitHub ansehen↗48
  • ailab-cvc/gpt4toolsAvatar von AILab-CVC

    AILab-CVC/GPT4Tools

    771Auf GitHub ansehen↗

    GPT4Tools is an intelligent system that can automatically decide, control, and utilize different visual foundation models, allowing the user to interact with images during a conversation.

    Python
    Auf GitHub ansehen↗771
  • evolvinglmms-lab/otterAvatar von EvolvingLMMs-Lab

    EvolvingLMMs-Lab/Otter

    3,331Auf GitHub ansehen↗

    Otter is a framework and toolkit for the pretraining, fine-tuning, and evaluation of vision-language models. It provides a pipeline for training large language models to process high-resolution images and video frames, integrating visual encoders with textual token spaces. The system is designed for multi-visual input processing, allowing models to interpret multiple images or video sequences within a single prompt. It supports multi-round conversation management to maintain context across interactions for detailed scene comprehension and visual reasoning. The framework covers a full develop

    Pythonartificial-inteligencechatgptdeep-learning
    Auf GitHub ansehen↗3,331
Alle 30 Alternativen zu ToolBench anzeigen→