How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success
The main features of corl-team/vl-dac are: Critic-Based Algorithms.
Projects with overlapping indexed features include: hitsz-tmg/veripo — [📄 Paper Link] [🤗 VerIPO-7B-v1.0]. lifan-yuan/implicitprm — Free Process Rewards without Process Labels. open-reasoner-zero/open-reasoner-zero — An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model. prime-rl/prime — ✨ Getting Started • 📖 Introduction 🔧 Usage • 📃 Evaluation • 🎈 Citation • 🌻 Acknowledgement • 📈 Star History. rookie-joe/autopsv — This repository contains the official implementation of AutoPSV: Automated Process-Supervised Verifier, accepted at…
An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
✨ Getting Started • 📖 Introduction 🔧 Usage • 📃 Evaluation • 🎈 Citation • 🌻 Acknowledgement • 📈 Star History