awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
PRIME-RL avatar

PRIME-RL/PRIME

0
View on GitHub↗
1,863 stars·114 forks·Python·Apache-2.0·7 vues

PRIME

✨ Getting Started • 📖 Introduction 🔧 Usage • 📃 Evaluation • 🎈 Citation • 🌻 Acknowledgement • 📈 Star History

Features

  • Critic-Based Algorithms - Process reinforcement learning using implicit reward signals.
  • Dense Reward Optimization - Process reinforcement using implicit rewards.
  • Policy Optimization - Process reinforcement learning using implicit reward signals.

Historique des stars

Graphique de l'historique des stars pour prime-rl/primeGraphique de l'historique des stars pour prime-rl/prime

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Start searching with AI

Questions fréquentes

Que fait prime-rl/prime ?

✨ Getting Started • 📖 Introduction 🔧 Usage • 📃 Evaluation • 🎈 Citation • 🌻 Acknowledgement • 📈 Star History

Quelles sont les fonctionnalités principales de prime-rl/prime ?

Les fonctionnalités principales de prime-rl/prime sont : Critic-Based Algorithms, Dense Reward Optimization, Policy Optimization.

Quelles sont les alternatives open-source à prime-rl/prime ?

Les alternatives open-source à prime-rl/prime incluent : thudm/treerl — Implementation for ACL'25 paper TreeRL: LLM Reinforcement Learning with On-Policy Tree Search. The implementation is… yangzhch6/treerpo — ```bash conda create -n rllm python=3.10 -y conda activate rllm. multimodal-art-projection/treepo — This is the official implementation of TreePO algorithm. ruc-nlpir/arpo — [ICLR 2026] Agentic Reinforced Policy Optimization (ARPO). trademaster-ntu/trademaster — TradeMaster is a reinforcement learning trading framework and algorithmic trading simulator designed for designing and… chenxinan-fdu/polaris — 🌠 A PO st-training recipe for scaling R L on A dvanced R eason I ng model S 🚀.

Alternatives open source à PRIME

Projets open source similaires, classés selon le nombre de fonctionnalités partagées avec PRIME.
  • ruc-nlpir/arpoAvatar de RUC-NLPIR

    RUC-NLPIR/ARPO

    1,049Voir sur GitHub↗

    ICLR 2026 Agentic Reinforced Policy Optimization (ARPO)

    Python
    Voir sur GitHub↗1,049
  • thudm/treerlAvatar de THUDM

    THUDM/TreeRL

    96Voir sur GitHub↗

    Implementation for ACL'25 paper TreeRL: LLM Reinforcement Learning with On-Policy Tree Search. The implementation is based on OpenRLHF

    Python
    Voir sur GitHub↗96
  • multimodal-art-projection/treepoAvatar de multimodal-art-projection

    multimodal-art-projection/TreePO

    65Voir sur GitHub↗

    This is the official implementation of TreePO algorithm.

    Python
    Voir sur GitHub↗65
  • yangzhch6/treerpoAvatar de yangzhch6

    yangzhch6/TreeRPO

    6Voir sur GitHub↗

    ```bash conda create -n rllm python=3.10 -y conda activate rllm

    Jupyter Notebook
    Voir sur GitHub↗6
Voir les 30 alternatives à PRIME→