awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
PRIME-RL avatar

PRIME-RL/PRIME

0
View on GitHub↗
1,863 نجوم·114 تفرعات·Python·Apache-2.0·6 مشاهدات

PRIME

✨ Getting Started • 📖 Introduction 🔧 Usage • 📃 Evaluation • 🎈 Citation • 🌻 Acknowledgement • 📈 Star History

Features

  • Critic-Based Algorithms - Process reinforcement learning using implicit reward signals.
  • Dense Reward Optimization - Process reinforcement using implicit rewards.
  • Policy Optimization - Process reinforcement learning using implicit reward signals.

سجل النجوم

مخطط تاريخ النجوم لـ prime-rl/primeمخطط تاريخ النجوم لـ prime-rl/prime

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

الأسئلة الشائعة

ما هي وظيفة prime-rl/prime؟

✨ Getting Started • 📖 Introduction 🔧 Usage • 📃 Evaluation • 🎈 Citation • 🌻 Acknowledgement • 📈 Star History

ما هي الميزات الرئيسية لـ prime-rl/prime؟

الميزات الرئيسية لـ prime-rl/prime هي: Critic-Based Algorithms, Dense Reward Optimization, Policy Optimization.

ما هي البدائل مفتوحة المصدر لـ prime-rl/prime؟

تشمل البدائل مفتوحة المصدر لـ prime-rl/prime: thudm/treerl — Implementation for ACL'25 paper TreeRL: LLM Reinforcement Learning with On-Policy Tree Search. The implementation is… yangzhch6/treerpo — ```bash conda create -n rllm python=3.10 -y conda activate rllm. multimodal-art-projection/treepo — This is the official implementation of TreePO algorithm. ruc-nlpir/arpo — [ICLR 2026] Agentic Reinforced Policy Optimization (ARPO). trademaster-ntu/trademaster — TradeMaster is a reinforcement learning trading framework and algorithmic trading simulator designed for designing and… chenxinan-fdu/polaris — 🌠 A PO st-training recipe for scaling R L on A dvanced R eason I ng model S 🚀.

بدائل مفتوحة المصدر لـ PRIME

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع PRIME.
  • ruc-nlpir/arpoالصورة الرمزية لـ RUC-NLPIR

    RUC-NLPIR/ARPO

    1,049عرض على GitHub↗

    ICLR 2026 Agentic Reinforced Policy Optimization (ARPO)

    Python
    عرض على GitHub↗1,049
  • thudm/treerlالصورة الرمزية لـ THUDM

    THUDM/TreeRL

    96عرض على GitHub↗

    Implementation for ACL'25 paper TreeRL: LLM Reinforcement Learning with On-Policy Tree Search. The implementation is based on OpenRLHF

    Python
    عرض على GitHub↗96
  • multimodal-art-projection/treepoالصورة الرمزية لـ multimodal-art-projection

    multimodal-art-projection/TreePO

    65عرض على GitHub↗

    This is the official implementation of TreePO algorithm.

    Python
    عرض على GitHub↗65
  • yangzhch6/treerpoالصورة الرمزية لـ yangzhch6

    yangzhch6/TreeRPO

    6عرض على GitHub↗

    ```bash conda create -n rllm python=3.10 -y conda activate rllm

    Jupyter Notebook
    عرض على GitHub↗6
عرض جميع البدائل الـ 30 لـ PRIME→