How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
This repository is a collection of specialized toolsets and libraries for large language model prompt engineering and security testing. It provides a library of advanced templates and frameworks designed to optimize the quality and specificity of model responses. The project includes resources for red teaming and security research, featuring a repository of prompts designed to bypass safety filters and operational constraints. It also provides techniques for system prompt extraction to reveal the internal instructions and configurations of AI personas. The collection covers a broader surface
A modular RL library to fine-tune language models to human preferences
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
The main features of pku-alignment/safe-rlhf are: Instruction Tuning Datasets, RLHF Frameworks, Jailbreak Defenses.
Projects with overlapping indexed features include: cyberalbsecop/awesome_gpt_super_prompting — This repository is a collection of specialized toolsets and libraries for large language model prompt engineering and… allenai/rl4lms — A modular RL library to fine-tune language models to human preferences. amadeuszhao/qmllm. anthropics/constitutionalharmlessnesspaper — This repository provides supplementary material for our paper Constitutional AI: Harmlessness from AI Feedback. carperai/trlx — trlx is a reinforcement learning library and training framework designed to align large language models using human… allenai/finegrainedrlhf — Fine-Grained RLHF.