How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
Implementation of the training framework proposed in Self-Rewarding Language Model , from MetaAI
This repo contains code and instructions for reproducing the experiments in the paper "RLCD: Reinforcement Learning from Contrast Distillation for Language Model Alignment" (https://arxiv.org/abs/2307.12950), by Kevin Yang, Dan Klein, Asli Celikyilmaz, Nanyun Peng, and Yuandong Tian. RLCD is a…
1. STaR 2. Mesh Transformer JAX 1. Updates 3. Pretrained Models 1. GPT-J-6B 1. Links 2. Acknowledgments 3. License 4. Model Details 5. Zero-Shot Evaluations 4. Architecture and Usage 1. Fine-tuning 2. JAX Dependency 5. TODO
The main features of ezelikman/star are: Self-Improvement Methods, Single Agent Optimization.
Projects with overlapping indexed features include: lucidrains/self-rewarding-lm-pytorch — Implementation of the training framework proposed in Self-Rewarding Language Model , from MetaAI. chengsong-huang/r-zero — Check out our paper or webpage for the details. facebookresearch/rlcd — This repo contains code and instructions for reproducing the experiments in the paper "RLCD: Reinforcement Learning… hkust-nlp/mstar — :star: Project Page . iamhankai/forest-of-thought — Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning. chengpengli1003/cort.