awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
PKU-Alignment avatar

PKU-Alignment/safe-rlhf

0
View on GitHub↗
1,605 stars·133 forks·Python·Apache-2.0·13 viewspku-beaver.github.io↗

Safe Rlhf

Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback

Features

  • Instruction Tuning Datasets - Contains safety-focused preference data for alignment training.
  • RLHF Frameworks - Framework for training value-aligned models with safety constraints.
  • Jailbreak Defenses - Implements safe reinforcement learning from human feedback.

Star history

Star history chart for pku-alignment/safe-rlhfStar history chart for pku-alignment/safe-rlhf

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Safe Rlhf

These projects share indexed features with Safe Rlhf. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • cyberalbsecop/awesome_gpt_super_promptingCyberAlbSecOP avatar

    CyberAlbSecOP/Awesome_GPT_Super_Prompting

    3,654View on GitHub↗

    This repository is a collection of specialized toolsets and libraries for large language model prompt engineering and security testing. It provides a library of advanced templates and frameworks designed to optimize the quality and specificity of model responses. The project includes resources for red teaming and security research, featuring a repository of prompts designed to bypass safety filters and operational constraints. It also provides techniques for system prompt extraction to reveal the internal instructions and configurations of AI personas. The collection covers a broader surface

    HTMLadversarial-machine-learningagentai
    View on GitHub↗3,654
  • allenai/rl4lmsallenai avatar

    allenai/RL4LMs

    2,390View on GitHub↗

    A modular RL library to fine-tune language models to human preferences

    Python
    View on GitHub↗2,390
  • amadeuszhao/qmllmA

    Amadeuszhao/QMLLM

    0View on GitHub↗
    View on GitHub↗0
  • allenai/finegrainedrlhfA

    allenai/FineGrainedRLHF

    0View on GitHub↗

    Fine-Grained RLHF

    View on GitHub↗0
Compare all 30 related projects→

Frequently asked questions

What does pku-alignment/safe-rlhf do?

Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback

What are the main features of pku-alignment/safe-rlhf?

The main features of pku-alignment/safe-rlhf are: Instruction Tuning Datasets, RLHF Frameworks, Jailbreak Defenses.

Which projects share features with pku-alignment/safe-rlhf?

Projects with overlapping indexed features include: cyberalbsecop/awesome_gpt_super_prompting — This repository is a collection of specialized toolsets and libraries for large language model prompt engineering and… allenai/rl4lms — A modular RL library to fine-tune language models to human preferences. amadeuszhao/qmllm. anthropics/constitutionalharmlessnesspaper — This repository provides supplementary material for our paper Constitutional AI: Harmlessness from AI Feedback. carperai/trlx — trlx is a reinforcement learning library and training framework designed to align large language models using human… allenai/finegrainedrlhf — Fine-Grained RLHF.