How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data
The main features of wantbook-book/serl are: Single Agent Optimization, Unsupervised Reward Methods.
Projects with overlapping indexed features include: chengsong-huang/r-zero — Check out our paper or webpage for the details. wangqinsi1/vision-zero — A domain-agnostic framework enabling VLM self-improvement through competitive visual games. hkust-nlp/mstar — :star: Project Page . gpoesia/minimo — This is the implementation of the following paper:. ezelikman/star — 1. STaR 2. Mesh Transformer JAX 1. Updates 3. Pretrained Models 1. GPT-J-6B 1. Links 2. Acknowledgments 3. License 4.… chengpengli1003/cort.
A domain-agnostic framework enabling VLM self-improvement through competitive visual games
1. STaR 2. Mesh Transformer JAX 1. Updates 3. Pretrained Models 1. GPT-J-6B 1. Links 2. Acknowledgments 3. License 4. Model Details 5. Zero-Shot Evaluations 4. Architecture and Usage 1. Fine-tuning 2. JAX Dependency 5. TODO