awesome-repositories.com
المدونة
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
OpenPipe avatar

OpenPipe/ART

0
View on GitHub↗
8,630 نجوم·721 تفرعات·Python·apache-2.0·10 مشاهداتart.openpipe.ai↗

ART

ART is a platform for agentic training, providing a reinforcement learning framework, training environment, and compute orchestrator. It enables the improvement of multi-step agent reasoning and tool usage through group relative policy optimization and a judge-based reward modeling system.

The project features tools for model distillation to transfer capabilities from large teacher models to smaller architectures, as well as a system for capturing execution trajectories to generate synthetic training data. It supports specialized training workflows including supervised fine-tuning for baseline establishment and the creation of reproducible task scenarios.

The infrastructure manages GPU compute resources via ephemeral environment provisioning and hybrid local-remote execution. It includes capabilities for trajectory-based data capture, model checkpoint management, and the routing of low-rank adaptations for inference.

The system provides observability through agent workflow scoring, compute cost monitoring, and training metric tracking.

Features

  • Reinforcement Learning Optimizers - Implements reinforcement learning optimizers, specifically group relative policy optimization, to refine agent reasoning and tool usage.
  • Agent Environments - Provides reproducible environments and standardized APIs for agents to practice complex tasks and tool usage.
  • Agentic Training Frameworks - Provides a platform for creating reproducible task scenarios and capturing execution trajectories for agent training.
  • Training Execution - Executes model training and token generation across local GPUs and managed autoscaling clusters.
  • Custom Evaluation Judges - Defines custom evaluation criteria to guide judge models in prioritizing or penalizing specific response characteristics.
  • Synthetic Dataset Generators - Generates synthetic training data by automatically creating diverse interaction tasks and task scenarios.
  • Group Relative Policy Optimization - Implements group relative policy optimization to stabilize reinforcement learning by comparing rewards across trajectory groups.
  • Reward Functions - Merges relative rankings from judge models with hand-crafted scores to create composite performance signals.
  • Training Loop Managers - Coordinates the iterative cycle of parallel inference rollouts and model weight updates.
  • Model Distillation Pipelines - Provides pipelines to transfer capabilities from large teacher models to smaller architectures via distillation.
  • Reinforcement Learning Reward Systems - Implements a judge-based reward modeling system that ranks agent trajectories to provide RL signals.
  • Model-Based Trajectory Ranking - Compares multiple agent execution paths and assigns rewards using a model-based judge.
  • Reinforcement Learning Training - Provides the execution engine for running training scenarios and updating model weights via reinforcement learning.
  • Group Relative Policy Optimization - Implements an iterative reinforcement learning loop using group relative policy optimization to refine agent reasoning.
  • RL Loop Integrations - Connects reinforcement learning loops to orchestration tools and server protocols to optimize multi-step reasoning.
  • Reward-Based Trajectory Management - Records sequences of system and user messages during rollouts to serve as data for reward-based optimization.
  • Trajectory-Based Agent Optimization - Refines model behavior by capturing and scoring sequences of tool calls and system messages.
  • Workflow Reward Scoring - Evaluates the correctness of multi-step agent workflows using reward functions to guide the training process.
  • GPU Training Clusters - Launches short-lived compute clusters for training tasks to decouple hardware management from training logic.
  • Training Orchestrators - Coordinates the transition between parallel inference rollouts and weight updates across distributed GPU hardware.
  • Agent Trajectory Logs - Logs agent interactions and tool calls during execution to generate training data for reinforcement learning.
  • Training Trajectory Capture - Logs sequences of system messages and tool calls as training examples for reinforcement learning and supervised fine-tuning.
  • LLM-As-A-Judge Scoring - Evaluates agent performance using a larger teacher model to rank trajectory quality against defined rubrics.
  • Protocol Training - Teaches models how to interact with MCP protocol servers to perform multi-step tool-based workflows.
  • Conversation Branching Systems - Stores multiple branching conversation histories within a single trajectory to support sub-agent interactions and delegations.
  • Non-Linear Trajectory Tracking - Stores multiple separate conversation histories within a single trajectory to support complex sub-agent interactions.
  • Synthetic Scenario Generators - Automatically generates diverse interaction tasks and edge cases to test external server integrations.
  • Knowledge Distillation - Provides tools for transferring capabilities from large teacher models to smaller, more efficient architectures.
  • Hybrid Execution Engines - Enables execution of inference and training across both local hardware and remote GPU backends.
  • Local Model Training Integrations - Runs training processes on user-owned hardware for a variety of open-weight model architectures.
  • Training Progress Monitoring - Tracks reward metrics and model performance over time to validate continuous improvement.
  • Model Performance Benchmarking - Benchmarks trained models against multiple baselines using validation sets to measure accuracy improvements.
  • Low-Rank Adaptation - Dynamically loads and serves low-rank adaptation weights to specialize agent behavior for different tasks.
  • Model Distillation Tools - Includes tools for transferring knowledge from large teacher models to smaller, efficient student architectures.
  • LoRA Adapter Interfaces - Provides a backend that dynamically loads and switches between LoRA adapters during inference to improve agent reliability.
  • Sub-Agent Trajectory Training - Supports training with multiple separate conversation histories within a single trajectory for sub-agent delegation.
  • Supervised Fine-Tuning - Supports supervised fine-tuning to establish model baselines and format adherence before reinforcement learning.
  • Adapter-Based Warm-Starts - Allows reinforcement learning to start from existing adapters to stabilize early training and reduce compute costs.
  • Synthetic Data Generators - Generates custom models without pre-labeled datasets by automatically producing inputs and evaluating performance.
  • Training - Automatically manages inference and training hardware to reduce operational overhead.
  • Ephemeral GPU Environments - Launches short-lived compute clusters to decouple hardware management from agent training logic.
  • Model Training Metrics - Logs critical training data including rewards, loss, and throughput to observability platforms.
  • Fine-Tuning Frameworks - Framework for training multi-step agents using GRPO.
  • Fine-Tuning Frameworks - Framework for training multi-step agents.

سجل النجوم

مخطط تاريخ النجوم لـ openpipe/artمخطط تاريخ النجوم لـ openpipe/art

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

بدائل مفتوحة المصدر لـ ART

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع ART.
  • oumi-ai/oumiالصورة الرمزية لـ oumi-ai

    oumi-ai/oumi

    8,858عرض على GitHub↗

    Oumi is a comprehensive large language model development platform designed for synthesizing data, fine-tuning models, and running performance evaluations. It serves as a unified environment for the entire model lifecycle, encompassing a training and fine-tuning suite, an evaluation framework, and tools for synthetic data generation and model distillation. The platform is distinguished by its iterative, failure-driven synthesis approach, which analyzes model weaknesses during evaluation to generate targeted training data. It utilizes an LLM-based judge framework to programmatically score respo

    Pythondpoevaluationfine-tuning
    عرض على GitHub↗8,858
  • zhaochenyang20/awesome-ml-sys-tutorialالصورة الرمزية لـ zhaochenyang20

    zhaochenyang20/Awesome-ML-SYS-Tutorial

    5,371عرض على GitHub↗

    This project provides a comprehensive technical guide and framework for engineering large-scale machine learning systems. It covers the full lifecycle of model development, focusing on the infrastructure and computational principles required to build, train, and serve generative AI models across distributed GPU clusters. The repository distinguishes itself by offering deep-dive tutorials and implementation strategies for complex system challenges. It emphasizes high-performance architectural primitives, such as collective communication orchestration, distributed tensor sharding, and static gr

    Python
    عرض على GitHub↗5,371
  • kiln-ai/kilnالصورة الرمزية لـ kiln-ai

    kiln-ai/kiln

    4,910عرض على GitHub↗

    Kiln is an LLM development workbench and evaluation framework designed for designing, testing, and optimizing prompts and AI agents. It functions as a multi-agent orchestrator and a RAG optimization tool, providing a visual interface for the iterative development of AI systems. The project distinguishes itself through a comprehensive fine-tuning pipeline that supports zero-code model training and reasoning distillation. It enables the creation of hierarchical multi-agent systems where specialized actors coordinate via tool calling, and it implements a Model Context Protocol server to expose t

    Python
    عرض على GitHub↗4,910
  • pwhiddy/pokemonredexperimentsالصورة الرمزية لـ PWhiddy

    PWhiddy/PokemonRedExperiments

    7,774عرض على GitHub↗

    This project is a game AI training framework designed to develop and monitor reinforcement learning agents within a legacy game environment. It functions as a training and monitoring system that optimizes autonomous agents to complete game objectives through exploration and reward-based learning. The framework includes tools for game memory mapping and real-time trajectory visualization. These capabilities translate raw game memory addresses into visual coordinates, allowing agent movements and session data to be streamed to a map for the analysis of navigation patterns and area exploration.

    Jupyter Notebook
    عرض على GitHub↗7,774
عرض جميع البدائل الـ 30 لـ ART→

الأسئلة الشائعة

ما هي وظيفة openpipe/art؟

ART is a platform for agentic training, providing a reinforcement learning framework, training environment, and compute orchestrator. It enables the improvement of multi-step agent reasoning and tool usage through group relative policy optimization and a judge-based reward modeling system.

ما هي الميزات الرئيسية لـ openpipe/art؟

الميزات الرئيسية لـ openpipe/art هي: Reinforcement Learning Optimizers, Agent Environments, Agentic Training Frameworks, Training Execution, Custom Evaluation Judges, Synthetic Dataset Generators, Group Relative Policy Optimization, Reward Functions.

ما هي البدائل مفتوحة المصدر لـ openpipe/art؟

تشمل البدائل مفتوحة المصدر لـ openpipe/art: oumi-ai/oumi — Oumi is a comprehensive large language model development platform designed for synthesizing data, fine-tuning models,… zhaochenyang20/awesome-ml-sys-tutorial — This project provides a comprehensive technical guide and framework for engineering large-scale machine learning… kiln-ai/kiln — Kiln is an LLM development workbench and evaluation framework designed for designing, testing, and optimizing prompts… pwhiddy/pokemonredexperiments — This project is a game AI training framework designed to develop and monitor reinforcement learning agents within a… verl-project/verl — This project is a distributed training infrastructure designed for aligning large language models through… vibrantlabsai/ragas — Ragas is an evaluation framework designed to measure the performance of retrieval-augmented generation pipelines and…