How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
//: # (![Hugging Face Collection(https://img.shields.io/badge/Models-fcd022?style=for-the-badge&logo=huggingface&logoColor=000)]())
π«SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation
π€ HF Models and Datasets Collection | π Arxiv Preprint
The main features of qingyangzhang/empo are: Unsupervised Reward Methods.
Projects with overlapping indexed features include: gpoesia/minimo β This is the implementation of the following paper:. insightllm/rl-without-gt β [//]: # ([![Hugging Face Collection](https://img.shields.io/badge/Models-fcd022?style=for-the-badge&logo=huggingfacβ¦ kelaxon/ssr-zero β π«SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation. leaplabthu/absolute-zero-reasoner β βοΈ Algorithm Flow β’ π Results β¨ Getting Started β’ ποΈ Training β’ π§ Usage β’ π Evaluation π Citation β’ π»β¦ lili-chen/self-questioning-lm β Self-Questioning Language Models. chengsong-huang/r-zero β Check out our paper or webpage for the details.