How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
This repo contains code and instructions for reproducing the experiments in the paper "RLCD: Reinforcement Learning from Contrast Distillation for Language Model Alignment" (https://arxiv.org/abs/2307.12950), by Kevin Yang, Dan Klein, Asli Celikyilmaz, Nanyun Peng, and Yuandong Tian. RLCD is a…
Dromedary: towards helpful, ethical and reliable LLMs.
This repository is the official code repository for our paper Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing.
1. STaR 2. Mesh Transformer JAX 1. Updates 3. Pretrained Models 1. GPT-J-6B 1. Links 2. Acknowledgments 3. License 4. Model Details 5. Zero-Shot Evaluations 4. Architecture and Usage 1. Fine-tuning 2. JAX Dependency 5. TODO
An unofficial implementation of Self-Alignment with Instruction Backtranslation .
The main features of spico197/humback are: Self-Improvement Methods.
Projects with overlapping indexed features include: ezelikman/star — 1. STaR 2. Mesh Transformer JAX 1. Updates 3. Pretrained Models 1. GPT-J-6B 1. Links 2. Acknowledgments 3. License 4.… facebookresearch/rlcd — This repo contains code and instructions for reproducing the experiments in the paper "RLCD: Reinforcement Learning… ibm/dromedary — Dromedary: towards helpful, ethical and reliable LLMs. jaehunjung1/impossible-distillation — This repository is the official code repository for our paper Impossible Distillation: from Low-Quality Model to… lucidrains/self-rewarding-lm-pytorch — Implementation of the training framework proposed in Self-Rewarding Language Model , from MetaAI. project-baize/baize-chatbot — Let ChatGPT teach your own chatbot in hours with a single GPU!