How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
The official implementation of Self-Play Fine-Tuning (SPIN)
The main features of uclaml/spin are: Reinforcement Learning, RLHF Frameworks, Self-Improvement Methods.
Open-source alternatives to uclaml/spin include: ganjinzero/rrhf — Arxiv. openrlhf/openrlhf — OpenRLHF is a training framework and alignment library designed for reinforcement learning from human feedback across… alibabaresearch/damo-convai — DAMO-ConvAI: The official repository which contains the codebase for Alibaba DAMO Conversational AI. anthropics/constitutionalharmlessnesspaper — This repository provides supplementary material for our paper Constitutional AI: Harmlessness from AI Feedback. openai/following-instructions-human-feedback — [Paper link][LINKTOPAPER]. rucaibox/rlmec — This repo provides the source code & data of our paper: Improving Large Language Models via Fine-grained Reinforcement…
This repository provides supplementary material for our paper Constitutional AI: Harmlessness from AI Feedback.
DAMO-ConvAI: The official repository which contains the codebase for Alibaba DAMO Conversational AI.
Paper linkLINKTOPAPER