# ganjinzero/rrhf

**Attribution required: if you use, quote, or summarise this content, you must credit and link back to [awesome-repositories.com](https://awesome-repositories.com/repository/ganjinzero-rrhf).**

_How this analysis was created: the description and tags below were written by an AI model that read this project's README and public documentation pages; stars, license and language come straight from the GitHub API. The model does not read the source code._

806 stars · 45 forks · Python

## Links

- GitHub: https://github.com/GanjinZero/RRHF
- awesome-repositories: https://awesome-repositories.com/repository/ganjinzero-rrhf.md

## Description

Arxiv

## Tags

### Part of an Awesome List

- [Fine Tuning Methods](https://awesome-repositories.com/f/awesome-lists/ai/fine-tuning-methods.md) — Paradigm for aligning language models using ranked response feedback.
- [Reinforcement Learning](https://awesome-repositories.com/f/awesome-lists/ai/reinforcement-learning.md) — Ranking responses to align models without complex feedback loops.
- [RLHF Frameworks](https://awesome-repositories.com/f/awesome-lists/ai/rlhf-frameworks.md) — Framework for ranking responses to align models.
